Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,725
datasets available to search
ShareScore release 0.9.0
Dataset results
4,725 results for “Normalization”
Reconstructing 10-km-resolution direct normal irradiance dataset through a hybrid algorithm
<p>The 41-year (1982-2022) daily DNI dataset (CHDNI) reconstructed in this study has been uploaded, and stored in netcdf format. The one-year dataset comprises daily DNI estimates for either 365 or 366 days, with individual data files separately organized by year. Each daily file is stored in mat format and labeled as "pred_xxxxxyymm," where ‘xxxx' denotes the year, ‘yy' represents the month, and ‘mm' stands for the day. The geographical scope of CHDNI dataset spans from 3°N to 54°N in latitude and from 72°E to 136°E in longitude. The mat matrix, encapsulating the data, is configured with dimensions of 361 rows and 641 columns, measured in W/m2.</p> <p>If you want to use the CHDNI dataset for related scientific research, please contact us (Email: WHU_wjy@whu.edu.cn).</p> <ul> <li>Wu J, Niu J, Qi Q, Gueymard CA, Wang L, Qin W, et al. Reconstructing 10-km-resolution direct normal irradiance dataset through a hybrid algorithm. Renewable and Sustainable Energy Reviews 2024; 204: 114805.</li> </ul>
Assessment of changes in circRNA expression based on transcripts of genes encoding ADAMTS proteins in patients with non-small cell lung carcinoma compared to normal tissue
Open the record for dataset details and reuse information.
TCIA full and low dose CT images and volumetric normal pancreas segmentations
<p>This publicly available CT dataset was downloaded from 50 portal venous phase abdominal CT scans (3mm-slice thickness) performed in 50 patients with liver metastases— from The Cancer Imaging Archive (TCIA).These CT scans were acquired with routine radiation dose levels per standard clinical protocols on SOMATOM Definition Flash CT scanner (Siemens Healthcare, Forchheim, Germany).Each study had been postprocessed to include a second reconstructed CT dataset at a simulated 25% radiation dose level. The mean (standard deviation [SD]) radiation dose for the full dose and the reduced dose CT datasets were 14.7 mSv (8 mSv) and 3.7 mSv (2 mSv), respectively. After a curation process by radiologists, 9 of 50 cases were excluded [chronic calcific pancreatitis (n = 1), postsurgical changes in pancreas (n = 2) and pancreatic lesions (n = 6)]. Volumetric pancreas segmentations were done on all the remaining 41 cases [16 males and 25 females, mean (SD) patient age: 60.8 years (14.7 yrs)] independently by the two radiologists (R1 and R2). Segmentations were first done on the full dose CT dataset and then repeated on the reduced dose CT dataset by both R1 and R2 after a gap of at least 24 hours to allow for memory extinction.</p>
HTP normalized counts via Deseq2
<p>This data was gathered from the Human Trisome project orginally and was processed in a way to try to normalize it. </p>
Supplementary Table S1. Combined analysis of variance containing the degrees of freedom (DF), mean squares (MS), P value (P val.), mean, coefficient of experimental variation (CEV%) and selective accuracy (SA) for the traits of luminosity (L*), chromaticity a* (a*), chromaticity b* (b*), grain length (length, mm), grain width (width, mm), grain thickness (thickness, mm), mass of 100 grains (Mass, g), normal grains (Ng, %), water absorption (absorption, %), cooking time (Ct, min:s), and concentrations of potassium (K, g kg-1 dry matter - DM), phosphorus (P, g kg-1 DM), calcium (Ca, g kg-1 DM), magnesium (Mg, g kg-1 DM), iron (Fe, mg kg-1 DM), zinc (Zn, mg kg-1 DM), and copper (Cu, mg kg-1 DM) obtained in 25 common bean cultivars evaluated in four experiments carried out from 2019 to 2021
<p><strong><span>Table S1.</span></strong><span> Combined analysis of variance.</span></p> <p><strong><span>Indirect selection for multiple technological and nutritional traits in common bean cultivars under different degrees of multicollinearity</span></strong></p> <p><strong><span>Bragantia, 2024.</span></strong></p>
mcrpc_wgbs_normal_samples_hg38
<p> HDF5-backed RangedSummarizedExperiment for WGBS Data (hg38 CpG sites) for 10 normal samples from the paper 'Zhao, Shuang G., et al. "The DNA methylation landscape of advanced prostate cancer." <em>Nature genetics</em> 52.8 (2020): 778-789.'. </p>
Effects of treatment by synthetic strigolactone (SL) on WUE and NUE in two tomato varieties under normal and combined-stress conditions
<p>The present work aims to analyse the influence of strigolactones (SL) on WUE and NUE under combined stress (drought and phosphate deprivation) in two tomato TOMRES varieties. One variety was chosen because it was found to be less sensitive to nitrogen starvation (# 250) by UNA, while the other (# 270) is more or less as sensitive as the wild type used in previous experiments (M82).</p>
Cell type labels for all clustering and normalization combinations compared for CODEX multiplexed imaging
<p>We performed CODEX (co-detection by indexing) multiplexed imaging on four sections of the human colon (ascending, transverse, descending, and sigmoid) using a panel of 47 oligonucleotide-barcoded antibodies. Subsequently images underwent standard CODEX image processing (tile stitching, drift compensation, cycle concatenation, background subtraction, deconvolution, and determination of best focal plane), and single cell segmentation. Output of this process was a dataframe of nearly 130,000 cells with fluorescence values quantified from each marker. We used this dataframe as input to 1 of the 5 normalization techniques of which we compared z, double-log(z), min/max, and arcsinh normalizations to the original unmodified dataset. We used these normalized dataframes as inputs for 4 unsupervised clustering algorithms: k-means, leiden, X-shift euclidian, and X-shift angular.</p> <p>From the clustering outputs, we then labeled the clusters that resulted for cells observed in the data producing 20 unique cell type labels. We also labeled cell types by hiearchical hand-gating data within cellengine (cellengine.com). We also created another gold standard for comparison by overclustering unormalized data with X-shift angular clustering. Finally, we created one last label as the major cell type call from each cell from all 21 cell type labels in the dataset. </p> <p>Consequently the dataset has individual cells segmented out in each row. Then there are columns for the X, Y position in pixels in the overall montage image of the dataset. There are also columns to indicate which region the data came from (4 total). The rest are labels generated by all the clustering and normalization techniques used in the manuscript and what were compared to each other. These also were the data that were used for neighborhood analysis for the last figure of the manuscript. These are provided at all four levels of cell type level granularity (from 7 cell types to 35 cell types). </p>
Normalized NMR integration values from the metabolomic analysis of Drosophila larvae extracts from 2 genotypes at 3 time points.
<p>We measured the metabolites related to energy production using 1H nuclear magnetic resonance spectroscopy (NMR). No alterations in the levels of carbohydrate stores or free amino acids were found between control and Sema1ai animals, corroborating the notion that the main metabolic changes are in the lipid metabolism. The exception is the glycolytic amino acid alanine (elevated in Sema1ai animals), confirming alterations in glycolysis. The levels of the ß-alanine amino acid are markedly reduced in 256 h AEL or 10.5-day-old Sema1ai animals, probably indicating muscle degeneration in the severely obese larvae that is consistent with the deteriorated state and reduced movement of the 10-day-old (256 hours) mutant larvae. Gluconeogenesis is stimulated by high lactate, and the concentration of lactate is higher in Sema1ai larvae than controls, though the difference is not statistically significant. Glycolysis is stimulated by glucose and inhibited by citrate, an early intermediate of the citric acid cycle. The increased citrate levels in the 10.5-day-old Sema1ai larvae suggest that glycolysis is lower at this age, consistent with the increased level of glucose in the severely obese larvae. The fact that both gluconeogenesis and glycolysis pathways are simultaneously enhanced in Sema1ai larvae support the hypothesis that the animals defecting in adiposity signaling are in a state of perceived energy insufficiency despite having sufficient energy stored.</p>
Replication Dataset for Improving Normalized Hurricane Damages
<p>This dataset can be used to replicate the results in Martinez (2020) Improving Normalized Hurricane Damages. https://doi.org/10.1038/s41893-020-0550-5</p> <p>The dataset was constructed using data from Weinkle et al. Nature Sustainability https://doi.org/10.1038/s41893-018-0165-2</p> <p>Building cost information was obtained from Robert Shiller: http://www.econ.yale.edu/~shiller/data/Fig3-1.xls</p>
Raw data for manuscript Semantic context can mask intelligibility declines at above-conversational speech levels in normal-hearing listeners
<p>Raw data for the manuscript in doc file. <br> Copied from the Matlab .m file. used for the analysis.</p> <p>To be updated.</p> <p>For details, contact me at mfer@health.sdu.dk</p>
Internal Normal Mode Analysis applied to RNA flexibility and conformational changes
<p>We investigated the capability of internal normal modes to reproduce RNA dynamics and predict observed RNA conformational changes, and, notably, those induced by the formation of RNA-protein and RNA-ligand complexes. Here, we extended our iNMA approach developed for proteins to study RNA molecules using a simplified rep- resentation of RNA structure and its potential energy. In this study, we considered three main data sets to investigate different aspects : i) one based on single-stranded RNA molecules for which all-atom MD simulations were computed; ii) one based on the available structures belonged to a specific Rfam family; iii) one based on the transition from unbound to bound RNA.</p> <p><strong>In each folder</strong></p> <p><em>modes.dat</em>: results obtained by iNMA (frequency and normal modes)</p> <p><em>das1.dat</em>: conversion from internal to cartesian normal modes</p> <p>Each file <em>name_enm.pdb</em> refers to a PDB structure with a CG representation (RNA three-bead model).</p> <p><strong>Dataset 1</strong>: d1.zip</p> <p>For the first dataset, we provide MD simulations converted into CG representation (RNA three-bead model), PCA analysis, the results obtained by iNMA for different values of distance cut-off <em>R</em><sub><em>c</em> </sub> and some scripts.</p> <p>Matlab and python scripts: </p> <p><em>analysis_pca.py</em>: to extract the different principal components</p> <p><em>analysis_PCA.m</em>: to compute overlap and cumative overlap in each folder</p> <p><em>analysis_complete_new.m</em>: to summarize the results</p> <p><strong>Dataset 2</strong>: d2.zip</p> <p>For this dataset, we provide the structure ensemble for Rfam family and the results obtained by iNMA for different values of distance cut-off <em>R</em><sub><em>c</em> </sub>and some scripts.</p> <p>PDB files:</p> <p><em>allensemble.pdb</em>: ensemble of PDB structures for a given Rfam family</p> <p><em>allensemble_enm.pdb</em>: ensemble of PDB structures for a given Rfam family converted to CG representation (RNA three-bead model)</p> <p><em>allensemble_enm_new.pdb</em>: ensemble of PDB structures for a given Rfam family with the same number of atoms for each model converted to CG representation (RNA three-bead model)</p> <p><em>model.pdb</em>: reference PDB structure</p> <p><em>model_enm.pdb</em>: reference PDB structure converted to CG representation (RNA three-bead model)</p> <p>Matlab script: </p> <p><em>pca_xray_anal.m</em>: PCA analysis, overlap, cumulative overlap, rmsip and plots</p> <p><strong>Dataset 3</strong>: d3.zip</p> <p>PDB structure:</p> <p><em>bound.pdb</em>: bound structure</p> <p><em>unbound.pdb</em>: unbound structure</p> <p><em>diff.dat</em>: difference between bound and unbound structure after superimposition </p> <p>RMSD<em>n </em>with n a number: the first column represents <span class="math-tex">\(\sqrt{\beta/2}\)</span></p> <p>Matlab script:</p> <p><em>rmsd_anal.m</em>: analysis best mode based on RMSD</p> <p><strong>Application to the CrPV-IRES</strong>: IRES.zip</p> <p>PDB structures:</p> <p> <em>IRES_cg.pdb</em>: Coarse-grain structure based on the PDB ID 5IT9</p> <p><em>b_end001_01_70.pdb</em>, <em>b_end001_01_80.pdb, b_end001_01_90.pdb</em>: Example of modified structures using the first lowest modes and different amplitudes <span class="math-tex">\(\beta\)</span></p> <p><em>b_end002_03_50.pdb</em>, <em>b_end002_03_60.pdb, b_end002_03_70.pdb</em>: Example of modified structures using the third lowest modes and different amplitudes <span class="math-tex">\(\beta\)</span></p>
SHIFT OF ONLINE CONSUMER PURCHASING BEHAVIOR IN VIETNAM DURING THE COVID-19 AND IN THE NEW NORMAL CONTEXT
<p>The study’s findings contribute to clarify shift of online consumer purchasing behavior and factors affecting the frequency of consumers’ online purchases in pandemic and after the pandemic’s demise, in order to help business to developeffective marketing strategies and enhance their presence in the e-commerce sector. </p>
16 - Normal Science versus Extraordinary Science
<p>Session #16 of the REINFORCE Critical & Scientific Thinking course, entitled, 'Normal Science versus Extraordinary Science', is given by Stefano Gattei, Professor of Science in Society, at the University of Trento, Italy.</p> <p>During the session, Prof. Gattei continues to explore the works of Popper and Kuhn and their application to different aspects of science, focussing particularly on the ideas of Puzzles and Anomalies and the ways in which crises in science bring about scientific revolutions and, thus, new Paradigms.</p>
15 - Normal Science as Puzzle Solving
<p>Session #15 of the REINFORCE Critical & Scientific Thinking course, entitled, 'Normal Science as Puzzle-Solving', is given by Stefano Gattei, Professor of Science in Society, at the University of Trento, Italy.</p> <p>During the session, Prof. Gattei covers Thomas Kuhn's ideas of Pre-Science and Normal Science, looking at the concepts of Paradigms and Puzzles, before going on to examine the ways in which the work of Ptolemy was still being used by Kepler a thousand years later.</p>
TARDIS configuration and emulator weights and training data for "1991T-Like Type Ia Supernovae as an Extension of the Normal Population"
<p>This dataset contains two archives of data related to the paper "1991T-Like Type Ia Supernovae as an Extension of the Normal Population"<br> <br> The first dataset, <a href="https://zenodo.org/api/files/de696fe0-3280-44f2-8ef4-975b92fad260/TARDIS_Emulator_Config.tar.gz">TARDIS_Emulator_Config.tar.gz </a>, contains the atomic data used to run TARDIS and a template configuration file from which samples are generated including the flags for the physics implementation used.</p> <p>The second dataset, InferenceScripts.tar.gz, contains the trained probabilistic neural network, the training/validation data (Under NNData), and scripts used to train the model and load and evaluate the model. Scripts that perform inference on spectra, as well as a folder of observed spectra (Under CorrectedSpectra), are included as well. A conda environment yaml file is included to rebuild the Python environment required to run all of the scripts. For questions please email John O'Brien.</p>
Apache Point Observatory Lunar Laser-ranging Operation (APOLLO) normal point data: 2006 through 2020
<p>Normal point data from the Apache Point Observatory Lunar Laser-ranging Operation (APOLLO) covering the 15-year span from April 2006 through the end of 2020. APOLLO measures the earth-moon separation by recording the round-trip travel time of photons from the Apache Point Observatory to five retro-reflector arrays on the moon. The APOLLO data set, combined with the 50-year archive of measurements from other lunar laser ranging (LLR) stations, can be used to probe fundamental physics such as gravity and Lorentz symmetry, as well as properties of the moon itself. These normal points have a median nightly accuracy of 1.7mm, which is an order of magnitude better than other LLR stations. Data is provided in both the "JPL" format (.np) and the Consolidated Laser Ranging Data Format (.crd)</p>
Comparing Finnish universities' publication profiles using multidimensional field-normalized indicators - dataset
<p>This dataset from the VIRTA Publication Information Service consists of the metadata of 241,575 publications of Finnish universities (publication years 2016–2021) merged from yearly datasets downloaded from <a href="https://wiki.eduuni.fi/display/cscvirtajtp/Vuositasoiset+Excel-tiedostot">https://wiki.eduuni.fi/display/cscvirtajtp/Vuositasoiset+Excel-tiedostot</a>.</p> <p>The dataset contains following information:</p> <ul> <li>Organisation: name of the university</li> <li>Publication year: the year of publication</li> <li>Subfield: one of 66 fields of science based on Statistics Finland field of science classification (in Finnish), see the classification in English: <a href="https://www2.stat.fi/en/luokitukset/tieteenala/">https://www2.stat.fi/en/luokitukset/tieteenala/</a></li> <li>Peer-reviewed: 1=peer-reviewed publications, 0=not peer-reviewed publications</li> <li>Science communication: 1=publications aimed at professional and general audiences, 0=peer-reviewed and not peer-reviewed publications aimed at academic audience.</li> <li>Bibliodiversity: 1=peer-reviewed book publications (chapters, monographs and edited volumes) and conference articles, 0=peer-reviewed journal articles.</li> <li>Multilingualism: share of peer-reviewed publications in languages other than English (Finnish, Swedish and other languages).</li> <li>Domestic publishing: 1=peer-reviewed publications in journals and books published in Finland, 0=peer-reviewed publications in journals and books published outside Finland.</li> <li>Domestic collaboration: 1=peer-reviewed publications with co-authors from more than one Finnish university, 0=peer-reviewed publications without co-authors from more than one Finnish university.</li> <li>International collaboration: 1=share of peer-reviewed publications with co-authors affiliated with foreign institutions, 0=share of peer-reviewed publications without co-authors affiliated with foreign institutions.</li> <li>Research performance: 1=peer-reviewed outputs in JUFO levels 2 (“leading”) and 3 (“top”) publication channels, 0=peer-reviewed outputs in JUFO levels 1 (“basic”) and 0 (“other”) publication channels.</li> <li>Open access: 1=peer-reviewed open access publications, including gold, hybrid and green OA, 0=peer-reviewed closed publications.</li> </ul>
O teórico Avedis Donabedian e as ações de Enfermagem no parto normal
<p>Produção para a disciplina Epistemologia nas Ciências do Cuidado em Saúde, do PACCS/EEAAC/UFF sobre o teórico Avedis Donabedian. O vídeo sintetiza as bases teóricas do pensamento de Donabedian e define os conceitos dos 7 pilares da qualidade descritos pelo autor.</p>
TCGA RNA-Seq normalized rsem data, TCGA clinical data and mutational signature profiles
<p>Normalized rsem RNA-Seq data for each of the 33 TCGA tumor types named as "<em>tumortype</em>rnaSeq.tar", a file containing TCGA clinical data ("cliDat_tcga_18.tar") and a text file with single base substitution mutational signatures for each sample from COSMIC (https://cancer.sanger.ac.uk/cosmic, “signatureProfileSample.txt.zip”).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.