Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4,694

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

4,694 results for “data analysis”

Learn how ShareScore rates datasets ↗
zenodo40/100

Synthetic Data for Neutrophil Analysis: Sets with regular shapes and Gaussian noise

<p><strong>Synthetic Datasets with regular shapes and Gaussian noise.</strong></p> <p><strong>Part of the PhagoSight neutrophil tracking and analysis package (Henry, et al., PLOS ONE, 2013):</strong></p> <p>&nbsp;</p> <p>https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0072636</p> <p>http://www.phagosight.org</p> <p>https://github.com/phagosight/phagosight</p> <p>&nbsp;</p> <p>A series of synthetic data sets that reproduce different behaviour characteristics of migrating neutrophils were generated in MATLAB. The data sets consisted of six artificial neutrophils that travelled along paths that presented different conditions of tortuosity, times to activation and proximity to other neutrophils during 98 time frames.</p> <p>Numerous data sets of neutrophils in zebrafish were carefully observed before setting the characteristics. Six trajectories were manually determined by setting the row, column positions of the centroids at every time point for 98 time frames. Each trajectory was designed so that it would represent different neutrophil behaviours: some trajectories were very oriented and had movements with uniform distance between time frames, whilst others were less uniform and would move at different velocities, some were tortuous whilst others were straight. The trajectories of cells 1 and 2 collided several times in the second half of the time frames whilst cells 3 and 4 collided at the beginning of the movement. Cell 6 migrated without meandering and then stopped at the end (which represents the wound area of an inflammation-based experiment) whilst 5 presented a delayed activation.&nbsp;</p> <p>Each time frame consisted of 11 slices of z-stack each with 275 x 275 pixels, where the neutrophils were formed by Gaussian distributions of higher intensities than the background and <strong>Gaussian noise </strong>(check the corresponding irregular shapes with Poisson noise plus another set with a <strong>single large neutrophil</strong> and Poisson noise). The orientation of the Gaussians varied according to the displacement of the artificial neutrophils,&nbsp;<em>i.e.</em>they were round when the cells were static, or elongated when in movement. The tracks with the Gaussians were saved as the&nbsp;<em>gold standard</em> and five different data sets were generated by adding varying levels of white Gaussian noise resulting in data sets with distributions with increasing similarity between the neutrophils and the background reflected by the decreasing values of the Bhattacharyya Distance (1.61, 1.25, 1, 0.66, 0.45) as defined by Coleman 1979.</p> <p>&nbsp;</p> <p>Files corresponding to the sets with irregular shapes and Poisson noise (noise increases from 1 to 6):</p> <ul> <li><strong>&nbsp;&nbsp;&nbsp; x,y,t trajectories &nbsp;&nbsp; ThreeDTracks</strong></li> <li><strong>&nbsp;&nbsp;&nbsp; Ground Truth&nbsp;&nbsp;&nbsp; &nbsp;&nbsp; syntheticData0_mat_Re </strong></li> <li><strong>&nbsp;&nbsp;&nbsp; First data set&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp; syntheticData1_mat_Re</strong></li> <li><strong>&nbsp;&nbsp;&nbsp; Second data set&nbsp;&nbsp;&nbsp;&nbsp; syntheticData2_mat_Re</strong></li> <li><strong>&nbsp;&nbsp;&nbsp; Third data set&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; syntheticData3_mat_Re</strong></li> <li><strong>&nbsp;&nbsp;&nbsp; Fourth data set&nbsp; &nbsp;&nbsp; syntheticData4_mat_Re</strong></li> <li><strong>&nbsp;&nbsp;&nbsp; Fifth data set&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; syntheticData5_mat_Re</strong></li> <li><strong>&nbsp;&nbsp;&nbsp; Sixth data set&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; syntheticData6_mat_Re</strong></li> </ul> <p>&nbsp;</p> <p>Corresponding GIF files are also included as illustrations of the cells in motion.</p> <p>&nbsp;</p> <p>Main Reference:</p> <p><a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0072636"><strong><em>PhagoSight</em>: An Open-Source MATLAB&reg; Package for the Analysis of Fluorescent Neutrophil and Macrophage Migration in a Zebrafish Model</strong> </a><br> Henry&nbsp;KM, Pase&nbsp;L, Ramos-Lopez&nbsp;CF, Lieschke&nbsp;GJ, Renshaw&nbsp;SA, Reyes-Aldasoro CC. (2013) <em>PhagoSight</em>: An Open-Source MATLAB&reg; Package for the Analysis of Fluorescent Neutrophil and Macrophage Migration in a Zebrafish Model. PLOS ONE 8(8): e72636. <a href="https://doi.org/10.1371/journal.pone.0072636">https://doi.org/10.1371/journal.pone.0072636</a></p>

opencc-by-4.0Apr 2013View details →
zenodo40/100

Data for the analysis from "Evidence for positive priming of leaf litter decomposition by contact with eutrophic pond sediments"

<p>These are the data files used in the analysis of the results of the experiments that are reported in the manuscript &quot;Evidence for positive priming of leaf litter decomposition by contact with eutrophic pond sediments&quot;.&nbsp; More details on the analysis can be found in at:&nbsp;https://github.com/KennyPeanuts/sediment_priming</p>

opencc-by-4.0May 2019View details →
zenodo40/100

Patients data for meta-analysis of genotype-phenotype associations in Bardet-Biedl Syndrome

<p>Data used for metaanalysis of the genotype-phenotype relationship in Bardet Biedl Syndrome.</p> <p>File &quot;EV table 1 literature.xlsx&quot; describes studies that were included in the metaanalysis. File &quot;EV table 2 dataset.xlsx&quot; contains individual patient data. Each row corresponds to a patient. If the same patient was reported in more than 1 study, their data were merged into one row. The columns are as follows:</p> <p>* source - a citation to the study the patient originated in</p> <p>* FamilyID - randomly generated ID of a family (unique over the dataset), two persons with the same FamilyID are related.</p> <p>* source case n. - A unique identifier of the patient within the study</p> <p>* gene - A gene carrying the principal BBSome related mutation</p> <p>* nucleotide change (allele 1,2)&nbsp; - description of the mutations in DNA individual alleles of the gene, in HGVS nomenclature</p> <p>* protein change (allele 1,2)&nbsp; - description of how the mutations in DNA change the resulting protein, in HGVS nomenclature</p> <p>* type of mut allele 1,2 - whether the given mutation&nbsp; is considered missense (MS) or large truncation (trunc)</p> <p>* mut/mut - combination of mutations for both alleles</p> <p>* additional mutations - mutations in other BBSome-related genes. Format is &quot;gene: DNA mutation, protein mutation&quot;</p> <p>* sex - &quot;F&quot; or &quot;M&quot;&nbsp; (where reported)</p> <p>* age group - age group (where reported)</p> <p>* age - age in years. Contains fractions, decimal values and &quot;5 month&quot;</p> <p>* RD, OBE, PD, CI, REP, REN, HEART, LIV, DD - presense or absence of phenotypes, if reported. RD &ndash; retinal dystrophy, OBE &ndash; obesity, PD &ndash; polydactyly, CI &ndash; cognitive impairment , REP &ndash; reproductive system anomalies, REN &ndash; renal anomalies, HRT &ndash; heart disease, LIV &ndash; liver anomalies, DD - Developmental delay. Values are &quot;&quot; (not reported), &quot;0&quot; (no phenotype), &quot;1&quot; (phenotype present), &quot;1!&quot; conflicting reports of phenotype in multiple studies (some patients were involved in multiple studies)</p> <p>* ethnicity - ethnicity of the patient, if reported</p> <p>* ethinc group - grouping of the ethnicities into 8 larger groups (see paper for details)</p> <p>* note - miscellanous text, in particular contains notes on patients merged from multiple studies</p> <p>====</p> <p>The protocol for this meta-analysis was pre-registered with PROSPERO (CRD42018096099).</p> <p>PubMed and Google Scholar databases were searched in May 2018 for the following keywords: [bardet-biedl syndrome AND (genotype phenotype OR cohort)]. Other suitable records were identified by snowball searching, in particular, by retrieving relevant articles from the references of the studied full-texts. In addition, all the references included in the publicly available Euro-Wabb database (<a href="https://lovd.euro-wabb.org/home.php">https://lovd.euro-wabb.org/home.php</a>) were covered. Our search was limited to the literature published in English language and covered the period from the inception of each database to the 21st of May 2018.</p>

opencc-by-sa-4.0Jan 2019View details →
zenodo40/100

Doctoral Thesis Artifact "User-Centered Tool Design for Data-Flow Analysis"

<p>This artifact contains the evaluation data and source code accompanying the doctoral thesis &quot;User-Centered Tool Design for Data-Flow Analysis&quot; by Lisa Nguyen Quang Do. The artifact contains (1) the survey questions and anonymized answers of the surveys conducted during the thesis, (2) the user study questionnaires, results, and test applications of the user studies conducted for the thesis, (3) the source code of the research prototypes and video demonstrations of their interfaces, and (4) the benchmark suites used for the empirical evaluation of those prototypes.</p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

A pooled analysis of the duration of chemoprophylaxis against malaria after treatment with artesunate-amodiaquine and artemether-lumefantrine: data & analysis output

<p>Data files and analysis output associated with the medRxiv preprint article: Bretscher et al.&nbsp;2019 A pooled analysis of the duration of chemoprophylaxis against malaria after treatment with artesunate-amodiaquine and artemether-lumefantrine.&nbsp;</p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

Proteomic data set of the analysis of black poplar (Populus nigra L.) seed storability

<p>Proteomic data set&nbsp;containing&nbsp;protein identification parameters (ESI MS/MS) and GO&nbsp;annotation functional classification (UniProt and QuickGO). Identification parameters of differentially abundant proteins of black poplar (<em>Populus nigra</em> L.) seeds stored in different temperature (3, -3, -20 and -196&deg;C) and time (12 and 24 months) conditions. Proteins were extracted and separated according to their isoelectric point (pI) and mass using 2-dimensional electrophoresis. Proteins that varied in abundance for temperature and time of storage were identified by mass spectrometry (ESI MS/MS). The mascot search algorithm (http://www.matrixscience.com) was used for protein identification against the NCBInr (http://www.ncbi.nig.gov) databases.Identified proteins were grouped due to biological process, molecular function and subcellular localization according to the gene ontology (GO) annotation using UniProt database and QuickGO search (https://www.ebi.ac.uk/QuickGO/).</p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

Bahamas National Hazard Analysis. Data Inputs and Outputs for the InVEST Coastal Vulnerability Model.

<p>The following folders contain the model inputs and outputs for the InVEST Coastal Vulnerability model that were used in the analysis discussed in:</p> <p>Silver JM, Arkema KK, Griffin RM, Lashley B, Lemay M, Maldonado S,<br> Moultrie SH, Ruckelshaus M, Schill S, Thomas A, Wyatt K and Verutes G<br> (2019) Advancing Coastal Risk Reduction Science and Implementation by<br> Accounting for Climate, Ecosystems, and People. Front. Mar. Sci. 6:556.<br> doi: 10.3389/fmars.2019.00556</p> <p>The readme.txt file contains information about data layers.</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

Details of offspring and source data for analysis of metabolic health and dietary preference in a rat model of acute alcohol exposure.

<p>This Excel file contains information on the number of offspring used to examine each outcome and the raw data for each data Table and Figure within a manuscript submitted to Journal of Physiology.&nbsp;</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

Data from: Analysis of leaf microbiome composition of near-isogenic maize lines differing in broad-spectrum disease resistance

<p>Data and code associated with the submitted manuscript &quot;Analysis of leaf microbiome composition of near-isogenic maize lines differing in broad-spectrum disease resistance&quot;. Detailed descriptions of each file can be found in the README.txt . The raw sequence data associated with this work can be downloaded from the NCBI Sequence Read Archive, listed under BioProject #PRJNA565009<strong>.</strong></p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

Data used in meta-analysis of debris-covered glacier melt

<p>The data used in the study Accuracy of Empirical Models of Debris-Covered Glaciers, A Winter-Billington, RD Moore and R Dadic, in prep. for submission to the Journal of Glaciology.</p>

opencc-by-4.0Oct 2019View details →
zenodo40/100

Bayesian network analysis of plasma microRNA sequencing data in patients with venous thrombosis

<p>This dataset contains the results of 2 related analyses, described in &quot;Bayesian network analysis of plasma microRNA sequencing data in patients with venous thrombosis&quot; (European Heart Journal Supplements, OUP). Link to the article: https://www.hal.inserm.fr/inserm-02310241</p> <p>1) In the directory &quot;miRNAs_MARTHA_GWAS&quot; : GWAS summary statistics for 162 circulating miRNAs in 344 VTE patients from the MARTHA cohort.</p> <p>Header for each summary file:</p> <p>Trait: miRNA id<br> chr: Chromosome<br> pos.hg19: Position of the variant in hg19/GRCh37 coordinates<br> SNP: rsid<br> A1: Reference allele on the forward strand<br> A2: Alternate allele on the forward strand<br> freq_A1: Frequency of reference allele<br> rsqr: Imputation quality defined by MACH<br> beta_A1: Estimated effect size (beta regression coefficient) of reference allele<br> se_A1: Estimated standard error of beta<br> p: p-value (significance of estimated beta)<br> z.score: Z-score</p> <p>&nbsp;</p> <p>2) In the directory &quot;meta_analysis&quot;: Random effect meta-analysis combining the results of our GWAS on the MARTHA cohort, and the results from a similar analysis conducted by Nikpay et al. (doi: 10.1093/cvr/cvz030). Summary statistics of 142 microRNAs, common to both datasets, were processed (and combine 1054 samples).</p> <p>Header for each summary file:</p> <p>chr: Chromosome<br> pos.hg19: Position of the variant in hg19/GRCh37 coordinates<br> SNP: rsid<br> A1: Reference allele on the forward strand<br> A2: Alternate allele on the forward strand<br> N: Sample size<br> Q: Cochran&#39;s heterogeneity statistic<br> Q.p: p-value of Cochran&#39;s Q<br> beta_A1: Estimated effect size (beta regression coefficient) of reference allele<br> se_A1: Estimated standard error of beta<br> p: p-value (significance of estimated beta)</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2018View details →
zenodo40/100

Data for Deines, Wang, & Lobell 2019 analysis on tillage and yields

<p>Data needed to reproduce the original analyses and figures from Deines et al. 2019 &quot;Satellites reveal a small positive yield effect from conservation tillage across the US Corn Belt&quot; (https://doi.org/10.1088/1748-9326/ab503b) can be found in the zipped `data` folder.</p> <p>This includes the field-level samples derived from the input annual tillage and yield map datasets, along with associated covariates. Due to privacy concerns, field locations are reported by county only (latitude and longitude have been removed), and the input map datasets are not publicly available but may obtained from their sources upon reasonable request.</p> <p>Unzipped file is ~11.3 gb; data is provided in a zipped folder to preserve file structure for best use with associated code repository located at <a href="https://doi.org/10.5281/zenodo.3525359">https://doi.org/10.5281/zenodo.3525359.</a></p>

opencc-by-3.0Oct 2019View details →
zenodo40/100

X-ray tomography image data of a graphite foam block (KFoam) and tortuosity analysis

<p>X-ray tomography (CT) image data of a graphite foam block (KFoam). The 3D image was generated with an X-ray tomography scan performed by Dr Llion Evans with Manchester X-ray Imaging Facility equipment, which was funded in part by the EPSRC (grants EP/F007906/1, EP/F001452/1 and EP/I02249X/1).</p> <p>The dataset includes: raw radiographs; scan &amp; reconstruction parameter settings file; reconstructed 3D volume. To visualise the 3D volume use software such as ImageJ (https://imagej.net/Fiji/Downloads). The volume image data (NMT_15_229_LLME_DivInterlayer.raw) is in binary format and has the following characteristics: 1586 x 1567 x 1588; 8-bit; little-endian byte order.</p> <p>The second .zip file is a 200 x 200 x 200 subset of this dataset. This was used to perform a tortuosity analysis on the foam. This dataset includes three sets of tiff images; tomographic slices; binarised slices; skeletonised slices. It also includes an excel file with the results of the tortuosity analysis performed with ImageJ.</p> <p>This data was used originally for the following publications (please cite if re-using the data):</p> <p>Ll.M. Evans, L. Margetts, P.D. Lee, C.A.M. Butler, E. Surrey, &ldquo;Image based in silico characterisation of the effective thermal properties of a graphite foam&rdquo;, Carbon, Vol. 143, pp. 542-558, 2018. <a href="https://doi.org/10.1016/j.carbon.2018.10.031">https://doi.org/10.1016/j.carbon.2018.10.031</a></p> <p>Ll.M. Evans, L. Margetts, P.D. Lee, C.A.M. Butler, E. Surrey, &ldquo;Improving modelling of complex geometries in novel materials using 3D imaging&rdquo;, Proceedings of NEA International Workshop on Structural Materials for Innovative Nuclear Systems, Manchester, UK, July 2016. <a href="https://www.oecd-nea.org/science/smins4/documents/P1-18_LlME_SMINS4_paper_reviewed.pdf">https://www.oecd-nea.org/science/smins4/documents/P1-18_LlME_SMINS4_paper_reviewed.pdf</a></p>

opencc-by-4.0Nov 2019View details →
zenodo40/100

epiGBS and RADseq data analysis

<p>This release contains the bismark coverage files and stacks pipeline output for analysis of methylation and DNA variation between 4 populations of oysters in Louisiana. This&nbsp;update&nbsp;adds the annotation reference files and a samples described file.</p>

opencc-by-4.0Nov 2019View details →
zenodo40/100

Text-fig. 1. D&E tree of Endress and Doyle (2009), from the combined morphological and molecular analysis of Doyle and Endress (2000), with modifications based on more recent data, showing the inferred evolution of the reticulum grading character (39). Boxes under names of taxa indicate their character state; shading of branches indicates their reconstructed state based on parsimony optimization with MacClade (Maddison and Maddison 2003). Nymph = Nymphaeales, Aust = Austrobaileyales, Chlor = Chloranthaceae, Piper = Piperales, Ca = Canellales, Magnol = Magnoliales. in Early Cretaceous Monocots: A Phylogenetic Evaluation

Text-fig. 1. D&amp;E tree of Endress and Doyle (2009), from the combined morphological and molecular analysis of Doyle and Endress (2000), with modifications based on more recent data, showing the inferred evolution of the reticulum grading character (39). Boxes under names of taxa indicate their character state; shading of branches indicates their reconstructed state based on parsimony optimization with MacClade (Maddison and Maddison 2003). Nymph = Nymphaeales, Aust = Austrobaileyales, Chlor = Chloranthaceae, Piper = Piperales, Ca = Canellales, Magnol = Magnoliales.

opencc-by-4.0Dec 2008View details →
zenodo40/100

Time-to-fatigue data for five Cypriniformes fish species and R script for data analysis

<p>The Excel file contains data from fixed velocity fatigue experiments for five small-sized Cypriniformes fish species. The recorded data includes common and scientific names of fish species, date and time of test trial, test flume length [cm], flow velocity treatment [cm/s], time-to-fatigue [sec], test water temperature [&deg;C], fish mass [g], fish fork length [cm], fish width [cm], and fish height [cm]. The readme text file explains the column names used in the Excel file. The Rscript file contains the code used to analyse the data.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Neutron activation analysis data of pottery from Cerro Mayal and the Chicama Valley, Peru

<p>Neutron activation analysis data of pottery from Cerro Mayal and other sites in the Chicama Valley, La Libertad Department, Peru.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Data from: Cross-sectional personal network analysis of adult smoking in rural areas

<p>This data package, titled&nbsp;<em>Data from: Cross-sectional personal network analysis of adult smoking in rural areas,</em> includes several files. First, there are annonymized raw data files in .rds file format (ego_data.rds &amp; alter_data.rds). Second, there is the R code that allow the replication of various statistical analyses. Interested parts may consult the R code as .pdf file format (Supplementary_Material_R_Code.pdf), .Rmd file format (that can be run to create the .pdf file format) and the .R file format (that can be accesed with R and RStudio). Moreover, the labels files are useful for recreating the Supplementary Material pdf file.&nbsp;</p> <p>Readers should know that this dataset corresponds to the study (paper)&nbsp;<em>Cross-sectional personal network analysis of adult smoking in rural areas.&nbsp;</em></p> <p>The ego_data.rds file includes 20 variables by 76 observations (respondents) while the alter_data.rds file includes 46 variables by 1681 observations (social contacts). We collected this information by deploying a personal network analysis research design. Initially, we interviewed 83 respondents (dubbed <em>egos</em>). Due to missing data, we kept in the analysis 76 egos and dropped seven respondents. We recruited the respondents using a link-tracining sampling framework. We started from a number of six seeds. We interviewed the seeds then we asked them to recommend other people in the study. We continued in a referee-referral fashion until 83 interviews were completed. The study was performed in a small rural Romanian community (4124 residents): Lerești (Argeș county).&nbsp;</p> <p>Our study was carried out in accordance with the recommendations, relevant guidelines, and regulations (specifically, those provided by the Romanian Sociologists Society, i.e., the professional association of Romanian sociologists). The research was performed in accordance with the Declaration of Helsinki. The research protocol was approved by a named institutional/licensing committee. Specifically, the Ethics Committee of the Center for Innovation in Medicine (InoMed) reviewed and approved all these study procedures (EC-INOMED Decision No. D001/09-06-2023 and No. D001/19-01-2024). All participants gave written informed consent. The privacy rights of the study participants were observed. The authors did not have access to information that could identify participants. Face-to-face interviews were collected between September 13 &ndash; 23, 2023, in Lerești, Romania. After each interview, information that could identity the participants were anonymized. Before conducting the interview, we provided each participant with a dossier containing informative materials about the project's objectives, how the data would be analyzed and reported, and their participation rights (e.g., the right to withdraw from the project at any time, even after the interview was completed). All study participants gave their written informed consent prior to enrolment in the study.</p> <p>The variables in the ego_data.rds file are as follows:</p> <p>(1) "networkCanvasEgoUUID" (unique alpha numeric code for each observation);&nbsp;</p> <p>(2) "ego_age" (the age of each study participant);&nbsp;</p> <p>(3) "ego_age.cen" (the age of each study participant, centered);&nbsp;</p> <p>(4) "ego_educ_b" (the education of each ego, binary);&nbsp;</p> <p>(5) "ego_educ_f" (the education of each ego, educational achievement);&nbsp;</p> <p>(6) "ego_marital.s_f" (the marital status of each ego);</p> <p>(7) "ego_occupation.cat2_f" (the occupation of each ego);&nbsp;</p> <p>(8) "ego_occupation_b" (the occupation of each ego, unemployed vs employed);&nbsp;</p> <p>(9) "ego_relstatus_b" (whether the ego is in a relationship or not);&nbsp;</p> <p>(10) "ego_sex_f" (the sex of the ego assigned at birth; male &amp; female);&nbsp;</p> <p>(11) "ego_sex_n" (the sex of the ego assigned at birth; 0 = male &amp; 1 = female);&nbsp;&nbsp;</p> <p>(12) "ego_smk_status_b1" (smoking status: 1 smoking, 0 others);</p> <p>(13) "ego_smk_status_b2" (smoking status: 1 former smoker, 0 others);&nbsp;&nbsp;</p> <p>(14) "ego_smk_status_b3" (smoking status: 1 not a smoker, 0 others);&nbsp;</p> <p>(15) "ego_smkstatus_f"&nbsp; (smoking status: former smoker, never-smoker, non-smoker (smoked too little), occasional smoker, smoker);&nbsp;</p> <p>(16) "ego_smoking_3cat"&nbsp; (smoking status: non-smoker, former smoker, smoker);</p> <p>(17) "net.size" (number of social contacts, alters, that were elicited by an ego);</p> <p>(18) "net.components" (number of strong components in the personal network);</p> <p>(19) "net.deg.centralization" (personal network degree centralization);</p> <p>(20) "net.density" (personal network density).&nbsp;</p> <p>The variables in the alter_data.rds file are as follows:</p> <p>(1) "alter_age" (the age of the alter);&nbsp;</p> <p>(2) "alter_age.cen" (the age of the alter - centered);&nbsp;</p> <p>(3) "alter_btw" (alter's betweenness score);&nbsp;</p> <p>(4) "alter_btw.cen" (alter's betweenness score - centered);&nbsp;</p> <p>(5) "alter_deg" (alter's degree score);&nbsp;</p> <p>(6)&nbsp; "alter_deg.cen" (alter's degree score - centered); &nbsp;</p> <p>(7) "alter_educ_b" (alter's education);&nbsp;</p> <p>(8) "alter_educ_f"&nbsp; (alter's education);&nbsp;</p> <p>(9) "alter_marital.s_f" (alter's marital status);&nbsp;</p> <p>(10) "alter_relstatus_b" (alter's marital status - binary variable);</p> <p>(11) "alter_sex_f" (alter's sex assigned at birth);</p> <p>(12) "alter_sex_n" (alter's sex assigned at birth; 1 - female; 0 - male);&nbsp;</p> <p>(13) "alter_smk_status_b1" (alter's smoking status; 1 smoker, 0 others);</p> <p>(14) "alter_smk_status_b2" (alter's smoking status; 1 former smoker, 0 others);</p> <p>(15) "alter_smk_status_b3" (alter's smoking status; 1 non-smoker, 0 others);</p> <p>(16) "alter_smoking_3cat" (alter's smoking status: three categories - smoker, non-smoker, former smoker);</p> <p>(17) "assortativity_score_fsmoker" (assortativity score for alter, former smoker);</p> <p>(18) "assortativity_score_nsmoker" (assortativity score for alter, non-smoker);</p> <p>(19) "assortativity_score_smoker" (assortativity score for alter, smoker);</p> <p>(20) "ego.alter_meet_f" (ego's meeting frequency with alter);&nbsp;</p> <p>(21) "ego_alter_meet_b" (ego's meeting frequency with alter, binary variable);</p> <p>(22) "ego.alter_meet_n" (ego's meeting frequency with alter, numerical codes);</p> <p>(23) "alter_rel.w.ego_f" (type of alters in an ego's network);</p> <p>(24) "networkCanvasUUID" (alpha numeric code for alter);</p> <p>(25) "networkCanvasEgoUUID" (alpha numeric code for ego);&nbsp;</p> <p>(26) "ego_smkstatus_f" (smoking status: former smoker, never-smoker, non-smoker (smoked too little), occasional smoker, smoker);&nbsp;</p> <p>(27) "ego_smoking_3cat" (three categories,&nbsp;smoking status: former smoker, non-smoker, smoker);</p> <p>(28) "ego_type_fsmk" (former smoking egos by type of ego-alter relationship);</p> <p>(29) "ego_type_nsmk" (non smoking egos by type of ego-alter relationship);</p> <p>(30) "ego_type_smk" (smoking egos by type of ego-alter relationship);</p> <p>(31) "ego_sex_f" (ego's sex, binary);</p> <p>(32) "ego_sex_n" (ego's sex, numerical code, 1 female, 0 male);&nbsp;</p> <p>(33) "ego_educ_b" (ego's education, binary variable)</p> <p>(34) "ego_age" (ego's age)</p> <p>(35) "ego_age.cen" (ego's age, centered)</p> <p>(36) "ego_relstatus_b" (ego's marital status, binary variable)</p> <p>(37) "ego_occupation_b" (ego's employment status, binary variable)</p> <p>(38) "net.components" (number of strong components in the personal network)</p> <p>(39) "net.deg.centralization" (degree centralization score in the personal networ)</p> <p>(40) "net.density" (density score in the personal network)</p> <p>(41) "prop_fsmokers" (proportion of former smokers in the personal network - alters)</p> <p>(42) "prop_fsmokers.cen" (proportion of former smokers in the personal network, centered- alters)</p> <p>(43) "prop_nsmokers" (proportion of non-smokers in the personal network- alters)</p> <p>(44) "prop_nsmokers.cen" (proportion of non-smokers in the personal network, centered- alters)</p> <p>(45) "prop_smokers" (proportion of smokers in the personal network- alters)</p> <p>(46) "prop_smokers.cen" (proportion of smokers in the personal network, centered- alters)</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

UniSpec: Deep Learning for Predicting the Full Range of Peptide Fragment Ion Series to Enhance the Proteomics Data Analysis Workflow

<p>UniSpec is a comprehensive DL spectrum predictor that can predict the intensity of the entire HCD MS/MS fragment ion series, going beyond existing tools limited to b/y ion series.&nbsp;</p> <p>All datasets developed for UniSpec model are shared on Zenodo as part of the UniSpec publication, "UniSpec: Deep Learning for Predicting Comprehensive Peptide Fragment Ion Series to Improve Peptide-Spectrum Matches from Shotgun Proteomics Experiments".</p> <p>This includes UniSpec datasets, downstream evaluation and analysis, and application case studies.</p> <p>1. pre-processed training, evaluation and testing data for machine learning;</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;UniSpec-Datasets.7z, Readme_UniSpecDatasets.txt</p> <p>2. Streamlined &nbsp;input datasets based on the fragmentation dictionary;</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; Streamlined_inputdatasets.7z, Readme_Streamlined_inputdatasets.txt</p> <p>3. Predictions on the validation and test sets;</p> <p>&nbsp; &nbsp; &nbsp; &nbsp;UniSpecPred_Validation-Test.7z, Readme_Predictons_ValidationTest.txt</p> <p>4. Evaluation by comparison with Prosit;</p> <p>&nbsp; &nbsp; &nbsp; a. Predictions: prosit_and_unispec_predictions.7z, Readme_prosit_and_unispec_predictions.txt</p> <p>&nbsp; &nbsp; &nbsp; b. Cosine similarity scores: prosit_vs_unispec_CS.7z, Readme_prosit_vs_unispec_CS.txt</p> <p>5. CSS for Different HCD Fragment Ion Series;</p> <p>&nbsp; &nbsp; &nbsp; &nbsp;CS_for_ion_splits.tsv</p> <p>6. Application 1: PSM rescoring;</p> <p>&nbsp; &nbsp; &nbsp; PSM rescoring_zipfiles.7z, &nbsp;PSM rescoring_readme.txt</p> <p>7. Application 2: In-silico spectral library search &nbsp;</p> <p>&nbsp; &nbsp; &nbsp; in-silico_librarysearch.7z, in-silico_librarysearch_readme.txt</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Can green hydrogen drive economic transformation in Saudi Arabia? - An input-output analysis of different Power-to-X configurations. Supplementary Data

<p>Supplementary material for peer review</p> <ul> <li>Modelling Data (input &amp; results)</li> <li>Literature Review</li> </ul>

opencc-by-4.0Sep 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record