Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,725
datasets available to search
ShareScore release 0.9.0
Dataset results
4,725 results for “Normalization”
Dataset of normalized 24-hour vectors
<p>Normalized data in 24-hour vectors representing pressure (P) and flow (Q) fluctuations during a day.</p> <p>dd_*.csv files are the pair-wise distance matrix for different measures (euclidean, DTW and GAK).</p> <p>anomalias_manuales.csv saves a boolean vector of manually selected anomalies during research.</p> <p> </p> <p>The proyect can be found at <a href="https://github.com/javialonsaso/TFM2020">github.com/javialonsaso/TFM2020.</a></p>
Normalized community CHP profiles
<p>Normalized community CHP profiles for Austria, France, Italy, Spain and Sweden for year 2018/2019.</p>
30-m Spatial Resolution Bioclimatic Dataset of 1991-2020 Climate Normals for Hubei Province, the Yangtze River Middle Reaches
<p><strong>Brief Introduction of the Dataset</strong></p> <p>This bioclimatic dataset is the product of research article "Mapping 30-m Resolution Bioclimatic Variables During 1991-2020 Climate Normals for Hubei Province, the Yangtze River Middle Reaches." published in IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing.</p> <p>The dataset contains 19 30-m resolution average bioclimatic variables during 1991-2020 Climate Normals for Hubei Province (108°21′42″—116°07′50″ E, 29°01′53″—33°6′47″ N), the core region of the Yangtze River middle reaches. The dataset was constructed by statistically downscaling the Climatic Research Unit (CRU) 1-km monthly climate variables (1440 in total), cablirating with ground observation data with 82 weather stations and aggregating based on the defination of 19 bioclimatic variables. The downscaling of four 1-km Climatic Research Unit monthly climate variables including monthly maximum, mean, minimum temperature and precipitation was firstly achieved by random forest model with 30-m resolution terrain and spatial data. Then the interpolation-based geographical differential analysis (GDA) was applied to improve the accuracy of downscaled products based on ground observation data. Finally, the bioclimatic variables were aggregated based on their definitions and averaged for the 30 years. The Yangtze River middle reaches is abundant of forestry, agriculture, biodiversity resources that requires finer bioclimatic data for better understands of these aspects. This dataset will provide higher spatial accuracy, more information and applicability in finer regional studies in the Yangtze River middle reaches.</p> <p> </p> <p><strong>Description of the 19 Bioclimatic Variables</strong></p> <p>The dataset contains 19 geotiff files in total. File names and the corresponding full name of bioclimatic variables are described as follows:</p> <p>Bio01 Mean annual air temperature (℃)<br>Bio02 Mean diurnal air temperature range (℃)<br>Bio03 Isothermality (%)<br>Bio04 Temperature seasonality (℃)<br>Bio05 Mean daily maximum air temperature of the warmest month (℃)<br>Bio06 Mean daily minimum air temperature of the coldest month (℃)<br>Bio07 Annual range of air temperature (℃)<br>Bio08 Mean daily mean air temperatures of the wettest quarter (℃)<br>Bio09 Mean daily mean air temperatures of the driest quarter (℃)<br>Bio10 Mean daily mean air temperatures of the warmest quarter (℃)<br>Bio11 Mean daily mean air temperatures of the coldest quarter (℃)<br>Bio12 Annual precipitation amount (mm)<br>Bio13 Precipitation amount of the wettest month (mm)<br>Bio14 Precipitation amount of the driest month (mm)<br>Bio15 Precipitation seasonality (%)<br>Bio16 Precipitation amount of the wettest quarter (mm)<br>Bio17 Precipitation amount of the driest quarter (mm)<br>Bio18 Precipitation amount of the warmest quarter (mm)<br>Bio19 Precipitation amount of the coldest quarter (mm)</p> <p> </p> <p><strong>Others</strong></p> <p>More information related to bioclimatic variables can be found on https://chelsa-climate.org/bioclim/</p>
BST/NOAA PSL Level 2 UAS Soil Moisture, Digital Elevation, Normalized Difference Vegetative Index, and Surface Temperature for SPLASH
<p>This dataset contains uncrewed aircraft systems (UAS) high-resolution data of soil moisture at the 0-5 cm soil depth, normalized difference vegetation index (NDVI), surface temperature, and digital elevation for the Study of Precipitation, the Lower Atmosphere, and Surface for Hydrology (SPLASH) campaign sponsored by the National Oceanic and Atmospheric Administration (NOAA). These data were collected near Avery Picnic (38.972425 degrees N,106.996855 degrees W) and Kettle Ponds (38.942005 degrees N,106.973006 degrees W) in the East River Watershed in Colorado from a series of flights starting on June 1st, 2022 and ending October 18th, 2023. Soil moisture measurements were retrieved using the Lobe Differencing Correlation Radiometer (LDCR) which is a L-Band (1-2 GHz) microwave radiometer and was flown on the E2 and S2 aerial platforms operated by Black Swift Technologies LLC. </p> <p> </p> <p>Each zip file contains a set of four Level 2 NetCDF files which provides the highest spatial resolution available for each of four products for a given flight location. With the Level 2 data, each flight location and variable can have different spatial resolutions depending on the sensor type, retrieval algorithm, and flight altitude. The file name convention for the zip files is as follows.</p> <p> </p> <p>uas_L2_yyyymmdd_hhmmss_vX.X.zip </p> <p>where</p> <p>L2 = Level 2 data </p> <p>yyyymmdd = year,month,day</p> <p>hhmmss = hour,minute,second</p> <p>vX.X = version number</p> <p>Time is the flight start time in UTC.</p> <p> </p> <p>The NetCDF file format contained in the zip files has a similar format to the zip files with convention</p> <p> </p> <p>uas_<var>_L2_yyyymmdd_hhmmss.nc </p> <p>where</p> <p><var> = vsm, dem, ndvi, or stmp</p> <p>vsm = volumetric soil moisture</p> <p>dem = digital elevation</p> <p>ndvi = normalized difference vegetation index</p> <p>stmp = surface temperature</p> <p> </p> <p>Note that each flight location using the E2 aerial platform required two flights so starting flight times for the soil moisture NetCDF files are different from the other three products.</p> <p><strong>November 2023 update</strong>: Version 2.0 added flight data from 2023. Version 2.0 includes an updated calibration of the soil moisture retrieval that has been applied to 2023 data, and a mask was applied to the soil moisture retrieval over water surfaces for both 2022 and 2023 data.</p> <p><strong>December 2023 update</strong>: Version 2.1 updated soil moisture data with a wet bias in v2.0 for flights #2 (17:40:35 UTC) and #3 (19:24:45 UTC) on July 27, 2022.</p>
Effects of Periodic Normal Stress Oscillations on Frictional Properties of Simulated Natural Fault Gouges under In Situ P-T Conditions
<p>Files named by in a format of "Uxxx_xx_xxMPa_xxC" refer to the original mechanical data recorded during experiment.</p> <p>The compressed package includes the files to perform numerical modeling, modeling results and the experimental data for comparison. To replicate the numerical modeling, readers can open the COMSOL project file (".mph" file) using COMSOL software (version >5.4) then input the parameters for the boundary conditions, such as the temperature, load-point velocity, oscillation amplitude and frequency. </p>
WikiMed and PubMedDS: Two large-scale datasets for medical concept extraction and normalization research
<p>Two large-scale, automatically-created datasets of medical concept mentions, linked to the <a href="https://uts.nlm.nih.gov/uts/umls/home">Unified Medical Language System (UMLS)</a>.</p> <p><strong>WikiMed</strong></p> <p>Derived from Wikipedia data. Mappings of Wikipedia page identifiers to UMLS Concept Unique Identifiers (CUIs) was extracted by crosswalking Wikipedia, Wikidata, Freebase, and the NCBI Taxonomy to reach existing mappings to UMLS CUIs. This created a 1:1 mapping of approximately 60,500 Wikipedia pages to UMLS CUIs. Links to these pages were then extracted as mentions of the corresponding UMLS CUIs.</p> <p>WikiMed contains:</p> <ul> <li>393,618 Wikipedia page texts</li> <li>1,067,083 mentions of medical concepts</li> <li>57,739 unique UMLS CUIs</li> </ul> <p>Manual evaluation of 100 random samples of WikiMed found 91% accuracy in the automatic annotations at the level of UMLS CUIs, and 95% accuracy in terms of semantic type.</p> <p><strong>PubMedDS</strong></p> <p>Derived from biomedical literature abstracts from <a href="https://pubmed.ncbi.nlm.nih.gov/">PubMed</a>. Mentions were automatically identified using distant supervision based on Medical Subject Heading (MeSH) headers assigned to the papers in PubMed, and recognition of medical concept mentions using the high-performance <a href="https://allenai.github.io/scispacy/">scispaCy</a> model. MeSH header codes are included as well as their mappings to UMLS CUIs.</p> <p>PubMedDS contains:</p> <ul> <li>13,197,430 abstract texts</li> <li>57,943,354 medical concept mentions</li> <li>44,881 unique UMLS CUIs</li> </ul> <p>Comparison with existing manually-annotated datasets (NCBI Disease Corpus, BioCDR, and MedMentions) found 75-90% precision in automatic annotations. Please note this dataset is <em>not </em>a comprehensive annotation of medical concept mentions in these abstracts (only mentions located through distant supervision from MeSH headers were included), but is intended as data for <em>concept n</em><em>ormalization</em> research.</p> <p>Due to its size, PubMedDS is distributed as 30 individual files of approximately 1.5 million mentions each.</p> <p><strong>Data format</strong></p> <p>Both datasets use JSON format with one document per line. Each document has the following structure:</p> <pre><code class="language-json">{ "_id": "A unique identifier of each document", "text": "Contains text over which mentions are ", "title": "Title of Wikipedia/PubMed Article", "split": "[Not in PubMedDS] Dataset split: <train/test/valid>", "mentions": [ { "mention": "Surface form of the mention", "start_offset": "Character offset indicating start of the mention", "end_offset": "Character offset indicating end of the mention", "link_id": "UMLS CUI. In case of multiple CUIs, they are concatenated using '|', i.e., CUI1|CUI2|..." }, {} ] }</code></pre> <p><strong>Version history</strong></p> <table align="left"> <thead> <tr> <th scope="col">Version</th> <th scope="col">Notes</th> </tr> </thead> <tbody> <tr> <td>1.0.0</td> <td>Initial release</td> </tr> <tr> <td>1.0.1</td> <td>Corrected duplication error in WikiMed.zip file</td> </tr> </tbody> </table> <p> </p> <p> </p> <p> </p> <p> </p>
COLONOMICS - predictive models for normal colon gene expression and DNA methylation for TWAS and MWAS
<p>We provide significant SNP prediction models derived from the COLONOMICS data (<a href="https://www.colonomics.org">https://www.colonomics.org</a>). Genotypes were obtained by Affymetrix 6.0 array, imputed to TopMed panel. Gene expression was obtained from Affymetrix U219 array, DNA methylation was obtained with Illuminan 450K array and miRNA expression was obtained by NGS. We provide SNP prediction models for 1,758 genes, 30,530 CpG probes and 38 miRNAs obtained from colon normal biopsy samples. These features can be predicted from SNPs located within ±1Mb, which we assumed they act through cis mechanisms. We include the model’s summary statistics and corresponding SNP weights in SQLite objects. Models were trained using the elastic net procedure employed in the PredictDB pipeline (<a href="https://predictdb.org/">https://predictdb.org</a>), according to which only models with a predictive performance p-value < 0.05 and R<sup>2</sup> > 0.1 are considered significant. We adjusted the models by basic covariates, i.e., sex, age, tissue type and colon anatomic location where biopsies were collected (left and right colon). Genome coordinates refer to GRCh37/hg19.</p>
Monthly Normalized MAVEN Data (2014-2021)
<p>These data were obtained from the magnetometer (MAG), Neutral Gas and Ion Mass Spectrometer (NGIMS), and Solar Wind Ion Analyzer (SWIA) onboard the Mars Atmosphere and Volatile EvolutioN (MAVEN) spacecraft. These two data sets were derived from these three instruments' data products and utilized in the article <strong>Influence of Magnetic Fields on Precipitating Solar Wind Hydrogen at Mars</strong>, which we plan to submit to Geophysical Research Letters.</p> <p>The backscattered and downward propagating penetrating proton data have been separated into two data files titled respectively. The contents of each file are as follows:</p> <p>1. Time [s]: epoch time (seconds since Jan. 1, 1970) for each penetrating proton measurement</p> <p>2. Orbit [#]: MAVEN's orbit number</p> <p>3. LSM [degrees]: Martian solar longitude</p> <p>4. Flux [eV/(eV s sr cm^2)]: peak angle averaged differential energy flux for each 4-s spectrum in electronvolt per electronvolt x second x steradian x square centimeters.</p> <p>5. Normalized Flux: monthly normalized peak flux (unitless)</p> <p>6. Energy [eV]: energy at which peak flux occurred for each 4-s spectrum found from Gaussian fit in electronvolts</p> <p>7. FWHM [eV]: full width at half maximum of 4-s spectrum found from Gaussian fit in electronvolts</p> <p>8. Altitude [km]: spacecraft altitude in kilometers </p> <p>9. SZA [degrees]: solar zenith angle in degrees </p> <p>10. X_MSO [km]: X position of MAVEN in the Mars-Sun-Orbit (MSO) coordinate system in kilometers</p> <p>11. Y_MSO [km]: Y position of MAVEN in the MSO coordinate system in kilometers</p> <p>12. Z_MSO [km]: Z position of MAVEN in the MSO coordinate system in kilometers</p> <p>13. Latitude [degrees]: Martian planetary latitude</p> <p>14. Longitude [degrees]: Martian planetary longitude</p> <p>15. Bx [nT]: X component of magnetic field measured by MAG in MSO coordinates in nanotesla </p> <p>16. By [nT]: Y component of magnetic field measured by MAG in MSO coordinates in nanotesla</p> <p>17. Bz [nT]: Z component of magnetic field measured by MAG in MSO coordinates in nanotesla</p> <p>18. Elevation Angle [degrees]: local magnetic field elevation angle measured relative to a vector tangent to the Martian surface during each penetrating proton measurement. 90 degree elevation angle corresponds to radial configurations while 0 degree elevation angle corresponds to horizontal configurations.</p> <p>19. Column Density [cm^-2]: CO<sub>2</sub> column density derived from NGIMS measurements in inverse square centimeters. All NaN values were used to fill for times during which NGIMS inbound verified data were not available, and the column density could not be determined.</p>
Python Time Normalized Superposed Epoch Analysis (SEAnorm) Example Data Set
<p>Solar Wind Omni and SAMPEX ( Solar Anomalous and Magnetospheric Particle Explorer) datasets used in examples for <a href="https://github.com/samwalton7645/SEA_Code">SEAnorm</a>, a time normalized superposed epoch analysis package in python.</p> <p>Both data sets are stored as either a HDF5 or a compressed csv file (csv.bz2) which contain a Pandas DataFrame of either the Solar Wind Omni and SAMPEX data sets. The data sets where written with pandas.DataFrame.to_hdf() and pandas.DataFrame.to_csv() using a compression level of 9. The DataFrames can be read using pandas.DataFrame.read_hdf( ) or pandas.DataFrame.read_csv( ) depending on the file format. </p> <p>The Solar Wind Omni data sets contains solar wind velocity (V) and dynamic pressure (P), the southward interplanetary magnetic field in Geocentric Solar Ecliptic System (GSE) coordinates (B_Z_GSE), the auroral electrojet index (AE), and the Sym-H index all at 1 minute cadence. </p> <p>The SAMPEX data set contains electron flux from the Proton/Electron Telescope (PET) at two energy channels 1.5-6.0 MeV (ELO) and 2.5-14 MeV (EHI) at an approximate 6 second cadence.</p> <p> </p>
Normal mode splitting function predictions for mantle anisotropy
<p>Predictions for normal mode splitting functions for 6 models of mantle anisotropy, accompanying the paper published in Geophysical Journal International by Restelli, Koelemeijer & Ferreira (2023). This is version 2 related to the revised manuscript. </p> <p>More details can be found in the README. </p>
Data to Three-Dimensional Binocular Eye-Hand Coordination in Normal Vision and with Simulated Visual Impairment
<p>This record contains experimental and analysis scripts (written in Matlab) as well as raw and processed data to reproduce the results shown in:</p> <p>Maiello, G., Kwon, M. & Bex, P.J. (2018) Three-dimensional binocular eye--hand coordination in normal vision and with simulated visual impairment. <em>Experimental Brain Research</em>. https://doi.org/10.1007/s00221-017-5160-8</p>
AllergyMap: An Open Source Corpus of Allergy Mention Normalizations
<p>AllergyMap is a mapping between free-text entered allergy medication to standard non-proprietary ontologies.</p>
Data on eye movements in people with glaucoma and peers with normal vision
<p>Eye movements were recorded from 44 elderly glaucoma patients and 32 age-similar healthy vision controls whilst watching three separate small video clips.</p>
Data from: Normalizing gas-chromatography–mass spectrometry data: method choice can alter biological inference
<p>Gas-Chromatography Mass Spectrometry data from European badger (<em>Meles meles</em>) sub-caudal gland secretion used in:</p> <p>Noonan, M.J., Tinnesand, H.V.,<sup> </sup>and Buesching, C.D. (2018). Normalizing gas-chromatography–mass spectrometry data: method choice can alter biological inference. BioEssays, 40(6): 0-0. DOI: 10.1002/bies.201700210.</p>
Herbarium specimen image of Asplenium normale, part of the collection of Finnish Museum of Natural History LUOMUS, University of Helsinki
Part of a training dataset of scanned herbarium specimens. The data paper and a summary landing page will be published on Zenodo as it gets published.<br><br>Content of this deposition:<br><br>- A JSON-LD datafile listing the label data associated with this herbarium specimen. The Darwin and Dublin Core data standards are used for most values.<br>- A JPEG image file of the scanned herbarium sheet.<br>- A lossless TIFF image from which the JPEG image has been derived.<br>- Two PNG files containing segmented image overlays of the scanned herbarium sheet. The _all extension indicates that all labels, color charts and pieces of text have received a different color against a black background color. The _sel extension indicates that these elements are white if they're barcode labels, yellow if they're color charts and red if they're anything else.
MERFISH data of the developing mouse visual cortex under normal- and dark-rearing
<p>MERFISH data for the manuscript "Spatial profiling of the interplay between cell type- and vision-dependent transcriptomic programs in the visual cortex"</p>
Database from: Orthometric, Normal and Geoid Heights In the Context of the Brazilian Altimetric Network
<p>This dataset is part of an article entitled "ORTHOMETRIC, NORMAL AND GEOID HEIGHTS IN THE CONTEXT OF THE BRAZILIAN ALTIMETRIC NETWORK" (https://doi.org/10.1590/s1982-21702022000100003), published in the Bulletin of Geodesic Sciences.</p> <p>This dataset includes 569 stations whose values for geodetic and normal height, gravity and geopotential numbers are provided by the IBGE (Brazilian Institute of Geography and Statistics). In addition, we included orthometric height data and differences between orthometric and normal height data calculated from the Hemlert and Mader methods, using constant and variable density data provided by Medeiros et al. (2021) (https://doi.org/10.1016/j.jsames.2021.103425) and Sheng et al. (2019) (https://doi.org/10.1016/j.tecto.2019.04.005).</p> <p> </p>
D.melanogaster Genelab OSD Normalized RNA Seq Matrix
<p><em>D.melanogaster </em>normalized counts RNA seq data matrix developed from NASA Genelab's open science data repository. Created using R.</p>
Transcriptomic analyses of normal-appearing CNS white matter from multiple sclerosis donors reveal subtype-specific molecular signatures of disease (REVISED)
<p>Datasets of bulk RNA-sequencing of NAWM from MS donors + supplementary images of RNA deconvolution of cell trajectories</p>
Data and scripts for reproducing "Optimisation and Analysis of Streamwise-Varying Wall-Normal Blowing in a Turbulent Boundary Layer"
<p>This is the accompanying data and Python scripts to reproduce the figures in "Optimisation and Analysis of Streamwise-Varying Wall-Normal Blowing in a Turbulent Boundary Layer", submitted to Flow, Turbulence and Combustion.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.