Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9,882
datasets available to search
ShareScore release 0.9.0
Dataset results
9,882 results for “scales”
Structure matters – Direct in-situ observation of cluster nucleation at atomic scale in a liquid phase (supplementary data)
<p>This a dataset of scanning transmission electron microscopy data showing Pt clusters nucleating in an ionic liquid. For each of the 4 movies there is the raw data (uncompressed .tif and compressed as .avi) and denoised versions (uncompressed .tif and compressed as .avi).</p> <p>This data is for the article "Structure matters – Direct in-situ observation of cluster nucleation at atomic scale in a liquid phase" published in ChemNanoMat (2020), by Trond R. Henninen, Debora Keller and Rolf Erni. (https://onlinelibrary.wiley.com/doi/full/10.1002/cnma.202000503)</p> <p><strong>Movie 1:</strong> Homogeneous nucleations of two clusters in a suspended thin film of ionic liquid. </p> <p><strong>Movie 2: </strong>Heterogeneous nucleation of a ca 8-9 atom cluster near the edge of a nanodroplet supported on a carbon film.</p> <p><strong>Movie 3: </strong>Heterogeneous nucleation of multiple clusters in a nanodroplet. Shortly after nucleation, they coalesce to form disordered nanoclusters.</p> <p><strong>Movie 4:</strong> Heterogeneous nucleation and dissolution cycles of spherical particles in a nanodroplet.</p>
Data from paper: "Large-scale variations in the dynamics of Amazon forest canopy gaps from airborne lidar data and opportunities for tree mortality estimates"
<p>Data from the paper:</p> <p>Dalagnol, R. <em>et al.</em> Large-scale variations in the dynamics of Amazon forest canopy gaps from airborne lidar data and opportunities for tree mortality estimates. <em>Sci Rep</em> <strong>11, </strong>1388 (2021). https://doi.org/10.1038/s41598-020-80809-w</p> <p>Link: https://www.nature.com/articles/s41598-020-80809-w</p> <p> </p> <p>This repository contains:</p> <p>1) Data frame with data from static and dynamic gaps used in Figure 2 (Dalagnol_2020_Data_Multitemporal_gaps.csv). Each row is the aggregated measurement at 5-km resolution. The site component referes to the five site studied with multitemporal data. Site order from 1 to 5 is DUC, TAP, FN1, BON and TAL.</p> <p>2) Data frame with data from static gaps and environmental factors used in Table 1, Figure 3, 4, 5 (Dalagnol_2020_Data_Singledate_gaps_Modeling.csv). Each row is the aggregated measurement of one site observed by airborne lidar data.</p> <p>3) Raster file at 5-km resolution with dynamic gap fraction estimates presented in Figure 5 (dynamic_gap_fraction_amazon.tif).</p> <p> </p> <p>If you need anything else, please contact the corresponding author: Ricardo Dalagnol (ricds@hotmail.com).</p>
Surface alkalinity, pH (total scale) and CO2 air-sea flux of the Mediterranean Sea under different alkalinisation scenarios.
<p>Surface maps and basin mean/total of annual mean surface alkalinity, pH (total scale) and CO2 air-sea flux of the Mediterranean Sea under different alkalinisation scenarios and for underlying the baseline projection (RCP4.5).</p> <p>Details on simulations and alkalinisation strategies are given in the reference article below.</p> <p> </p> <p>Reference:</p> <p>Butenschön, M., Lovato, T., Masina, S., Caserini, S., Grosso, M., 2021. Alkalinization Scenarios in the Mediterranean Sea for Efficient Removal of Atmospheric CO2 and the Mitigation of Ocean Acidification. Front. Clim. 3. <a href="https://doi.org/10.3389/fclim.2021.614537">https://doi.org/10.3389/fclim.2021.614537</a></p>
Dataset of Measurement and conceptualization of maternal PTSD following childbirth: Psychometric properties of the City Birth Trauma Scale – French version (City BiTS-F)
<p>The City Birth Trauma Scale (City BiTS-F) was developed to assess posttraumatic stress disorder following childbirth (PTSD-FC), based on the PTSD criteria of the DSM-5. Recent studies investigating the latent factor structure of PTSD-FC symptoms in women reported mixed results. Given that no validated French questionnaire exists to measure PTSD-FC symptoms, this study first aimed to validate the French version of the CBTS (City BiTS-F). Second, it aims to establish the latent factor structure of PTSD-FC.</p> <p>This dataset contains data on the mental health (i.e., PTSD-CB, depression, anxiety) of 541 mothers who gave birth during the last 12 months. Sociodemegraphic data such as maternal age, marital status, educational level, parity, gravidity, weeks of gestation, type of delivery, history of traumatic childbirth, or history of traumatic event is available. </p> <p>This dataset is related to: Sandoz, V., Hingray, C., Stuijfzand, S., Lacroix, A., El Hage, W., & Horsch, A. (2022). Measurement and conceptualization of maternal PTSD following childbirth: Psychometric properties of the City Birth Trauma Scale—French Version (City BiTS-F). <em>Psychological Trauma: Theory, Research, Practice, and Policy, 14</em>(4), 696–704. <a href="https://psycnet.apa.org/doi/10.1037/tra0001068">https://doi.org/10.1037/tra0001068</a></p>
Accessibility Indicators to services at EU scale - 1km grid indicators
<p>This archive makes available <strong>accessibility indicators at EU scale from populated 1km EU grid to towns and cities at EU scale</strong> (512 million travel time by car calculated between origins and destinations). It follows a reproducible, transparent and updatable framework. It uses <strong>only open source and free routing engines (OSRM)</strong>, based on OpenStreetMap (OSM) network. This routing engine makes possible the creation of travel time indicators for a large set of origins and destinations.</p> <p>The EU towns and cities layer has been recently made available and named by the European Commission. This layer is based on a <a href="https://ec.europa.eu/regional_policy/information-sources/maps/urban-centres-towns_en">common methodology</a> for all Europe. Within GRANULAR activities, we consider the towns and cities layer as <strong>a proxy</strong> to discuss on little and medium commercial centralities in Europe.</p> <p>This methodological framework, <strong>implemented with open source solutions (data and code) only and documented in a reproducible way in R notebooks</strong>, could be easily extended to other origins and destinations, if a relevant layer will be identified in the future.</p> <p>Based on travel time matrix, it is possible to compute a large set of indicators. This archive (see readme at the root folder) <strong>describes the input data used, summarises the data processing and provide information and metadata on output indicators created at 1km grid cells.</strong></p> <p>All the output data is also available. </p> <p> </p>
The role of injection method on residual trapping at the pore-scale in continuum-scale samples: segmented data
<p>The experiments in this work explore the role of a variable injection rate on gas saturation and residual trapping. There are 2 experiments in this work H2L (high to low injection rate) and L2H (low to high injection rate). The workflow for processing the micro-CT images to get the segmented images is described in [1]. </p><p>The following scans are included in this repository NB. all data for this repository is segmented micro-CT data.: </p><ol><li>Dry scan prior to experiment = merged_binning_2_38_1927</li><li>H2L during high flow = merged_segmented_flow_09_h2lh_merged</li><li>H2L during low flow = merged_segmented_flow_11_h2ll_2_merged</li><li>H2L at the end of drainage (no flow) =merged_segmented_flow_16_dra1_pd5_merged</li><li>H2L at the end of imbibition (no flow) =merged_segmented_flow_21_imb1_pi1_merged</li><li>L2H during low flow = merged_segmented_flow_29_2_l2hl_merged</li><li>L2H during high flow = merged_segmented_flow_30_l2hh_merged</li><li>L2H at the end of drainage (no flow) =merged_segmented_flow_31_dra2_pd1_merged</li><li>L2H at the end of imbibition (no flow) =merged_segmented_flow_33_imb2_pi1_merged</li></ol>
Exploring Large-Scale Entanglement in Quantum Simulation
<p>Here we provide data for the manuscript " <a href="https://arxiv.org/abs/2306.00057">Exploring Large-Scale Entanglement in Quantum Simulation</a> " with arXiv id <a href="https://arxiv.org/abs/2306.00057">"arXiv:2306.00057</a>". The data set contains both raw and analyzed data saved as ".mat files" Please see the uploaded readme file to understand the data structure. The peer-reviewed article will appear in the future. Please check the published article for recent figures. </p>
Centre frequencies and uncertainties for "Evidence for a kilometre-scale seismically slow layer atop the core-mantle boundary from normal modes"
<p>A table containing the centre frequencies and uncertainties used for the study presented in "Evidence for a kilometre-scale seismically slow layer atop the core-mantle boundary from normal modes". This table is the same as is contained in the supplementary materials of that paper.</p> <p>Russell, S., Irving, J. C. E., Jagt, L., & Cottaar, S. (2023). Evidence for a kilometer-scale seismically slow layer atop the core-mantle boundary from normal modes. Geophysical Research Letters, 50, e2023GL105684. <a href="https://doi.org/10.1029/2023GL105684">https://doi.org/10.1029/2023GL105684</a></p>
Global 1km Land Surface Parameters for Kilometer-Scale Earth System Modeling (SAI_2016_2020)
<p>Earth system models (ESMs) are progressively advancing towards the kilometer scale (k-scale). However, the surface parameters for Land Surface Models (LSMs) within ESMs running at the k-scale are typically derived from coarse resolution and outdated datasets. This study aims to develop a new set of global land surface parameters with a resolution of 1 km for multiple years from 2001 to 2020, utilizing the latest and most accurate available datasets. Specifically, the datasets consist of parameters related to land use and land cover, vegetation, soil, and topography. Differences between the newly developed 1k land surface parameters and conventional parameters emphasize their potential for higher accuracy due to the incorporation of the most advanced and latest data sources. To demonstrate the capability of these new parameters, we conducted 1 km resolution simulations using the E3SM Land Model version 2 (ELM2) over the contiguous United States. Our results demonstrate that land surface parameters contribute to significant spatial heterogeneity in ELM2 simulations of soil moisture, latent heat, emitted longwave radiation, and absorbed shortwave radiation. On average, about 31% to 54% of spatial information is lost by upscaling the 1 km ELM2 simulations to a 12 km resolution. Using eXplainable Machine Learning (XML) methods, the influential factors driving the spatial variability and spatial information loss of ELM2 simulations were identified, highlighting the substantial impact of the spatial variability and information loss of various land surface parameters, as well as the mean climate conditions. The comparison against four benchmark datasets indicates that ELM generally performs well in simulating soil moisture and surface energy fluxes. The new land surface parameters are tailored to meet the emerging needs of k-scale LSMs and ESMs modeling with significant implications for advancing our understanding of water, carbon, and energy cycles under global change.</p> <p>This data repository is linked to <a href="../records/10815170" target="_blank" rel="noopener">https://zenodo.org/records/10815170</a></p>
Global 1km Land Surface Parameters for Kilometer-Scale Earth System Modeling
<p><strong>Summary</strong>: Earth system models (ESMs) are progressively advancing towards the kilometer scale (k-scale). However, the surface parameters for Land Surface Models (LSMs) within ESMs running at the k-scale are typically derived from coarse resolution and outdated datasets. This study aims to develop a new set of global land surface parameters with a resolution of 1 km for multiple years from 2001 to 2020, utilizing the latest and most accurate available datasets. Specifically, the datasets consist of parameters related to land use and land cover, vegetation, soil, and topography. Differences between the newly developed 1k land surface parameters and conventional parameters emphasize their potential for higher accuracy due to the incorporation of the most advanced and latest data sources. To demonstrate the capability of these new parameters, we conducted 1 km resolution simulations using the E3SM Land Model version 2 (ELM2) over the contiguous United States. Our results demonstrate that land surface parameters contribute to significant spatial heterogeneity in ELM2 simulations of soil moisture, latent heat, emitted longwave radiation, and absorbed shortwave radiation. On average, about 31% to 54% of spatial information is lost by upscaling the 1 km ELM2 simulations to a 12 km resolution. Using eXplainable Machine Learning (XML) methods, the influential factors driving the spatial variability and spatial information loss of ELM2 simulations were identified, highlighting the substantial impact of the spatial variability and information loss of various land surface parameters, as well as the mean climate conditions. The comparison against four benchmark datasets indicates that ELM generally performs well in simulating soil moisture and surface energy fluxes. The new land surface parameters are tailored to meet the emerging needs of k-scale LSMs and ESMs modeling with significant implications for advancing our understanding of water, carbon, and energy cycles under global change.</p> <p><br><strong>Format</strong>: NetCDF.<br><strong>Institution</strong>: Atmospheric, Climate, and Earth Sciences Division, Pacific Northwest National Laboratory<br><strong>Contacts</strong>: Lingcheng Li (lingcheng.li@pnnl.gov; lingchengliwhu@gmail.com), Gautam Bisht (gautam.bisht@pnnl.gov)</p> <p><strong>Description</strong>: This dataset provides land surface parameters specifically designed for global kilometer scale earth system modeling.<br><strong>Spatial resolution</strong>: ~1 km, corresponding to 1/120 degree.<br><strong>Temporal resolution</strong>: includes yearly (2001-2020), monthly (2001-2020), and static data for different parameters.</p> <p><br><strong>Reference</strong>: <strong>Li, L., Bisht, G., Hao, D., and Leung, L.-Y. R.: Global 1km Land Surface Parameters for Kilometer-Scale Earth System Modeling, Earth Syst. Sci. Data Discuss. [preprint], https://doi.org/10.5194/essd-2023-242, Acceptance, 2023.</strong></p> <p>It includes four categories of parameters, Please refer to the readme file for details:<br>1. LULC: land use and land cover parameters<br>2. VEGE: vegetation paramertes<br>3. SOIL: soil parameters<br>4. TOPO: topography parameters</p> <p>Due to storage limitations, the LAI and SAI files are stored in the following repositories:</p> <p>1) LAI 2001-2005: <a href="../records/10815637" target="_blank" rel="noopener">https://zenodo.org/records/10815637</a>; 2) LAI 2006-2010: <a href="../records/10815649" target="_blank" rel="noopener">https://zenodo.org/records/10815649</a>; 3) LAI 2011-2015: <a href="../records/10815658" target="_blank" rel="noopener">https://zenodo.org/records/10815658</a>; 4) LAI 2016-2020: <a href="../records/10815662" target="_blank" rel="noopener">https://zenodo.org/records/10815662</a>;</p> <p>5) SAI 2001-2005: <a href="../records/10815623" target="_blank" rel="noopener">https://zenodo.org/records/10815623</a>; 6) SAI 2006-2010: <a href="../records/10815629" target="_blank" rel="noopener">https://zenodo.org/records/10815629</a>; 7) SAI 2011-2015: <a href="../records/10790724" target="_blank" rel="noopener">https://zenodo.org/records/10790724</a>; 8) SAI 2016-2020: <a href="../records/10790758" target="_blank" rel="noopener">https://zenodo.org/records/10790758</a></p>
Supporting data "Scaling theory for the statistics of slip at frictional interfaces"
<p>Principle data supporting "Scaling theory for the statistics of slip at frictional interfaces"</p> <p>T. W. J. de Geus and M. Wyart (2022), <em>Phys. Rev. E</em>, 106(6):065001.</p> <ul> <li>See code at <a href="../doi/10.5281/zenodo.10723197">doi: 10.5281/zenodo.10723197</a> (and its documentation) for workflow, detailed information of the data, and further dependencies. </li> <li>The files <code>N=*_Run*.zip</code> contain fully restorable events for event-driven athermal quasi-static shear. Sequentually numbered files contain different parts of a single dataset.</li> <li>The file <code>summary.zip</code> contains an extract of the key variables of these runs, and of triggers at different stresses. Finally, it contains "flow" data acquired by driving at finite rate. </li> <li>The files <code>N=3^6x4_Trigger_EnsemblePack.zip</code> contain fully restorable triggers at different stresses in the largest system. The sequentially numbered files correspond to one dataset split in different<em> </em><code>.h5</code> files.</li> <li>Highly specific (and poorly documentated) plotting functions are available upon request.</li> </ul>
Hydroelastic response of the scaled model of a floating offshore wind turbine platform in waves: HELOFOW Project Database
<p>This dataset contains the data measured during the <strong>HELOFOW </strong>model test campaign, performed at the Ocean and Hydrodynamic Engineering wave tank of Ecole Centrale Nantes (ECN): decay tests, regular wave tests and irregular waves tests. The preprocessed measured data is contained in MAT files.</p> <p>The model, the measurements and the tests are described in the appended Excel files. A Matlab(R) function is given as a short example to show how the MAT files are structured and how data may be handled for a plot. </p> <p>As stated in the reference paper (Leroy et al., <em>Ocean Engineering</em>, 2022):</p> <p>"As the size of floating wind turbines continues to increase, floating platforms reach dimensions that make their elastic and hydro-elastic behaviour significant. Several works in connection with the numerical modelling of the elastic behaviour of these wind turbines have been carried out but few validation data are available. This study focuses on the hydro-elastic response of a large floating wind turbine, in regular waves and severe sea-states. A new experimental wind turbine model has been designed to represent a 1:40 Froude-scaled spar platform carrying the DTU 10 MW turbine. The main challenge is here to reproduce a 1st bending mode frequency and hydrodynamic loads representative of a realistic large floating wind turbine. The platform model is made of a flexible backbone, reproducing the correct flexibility, and light floaters fixed on it provide the correctly scaled geometry. This experimental model is tested in various conditions including regular waves of several periods and steepness, and irregular waves of various intensity, including extreme 50-year return period conditions."</p> <p> </p> <p>This work was carried out within the framework of the WEAMEC, West Atlantic Marine Energy Community, and with funding from the Pays de la Loire Region and Europe (European Regional Development Fund). <br><br>HELOFOW project on <a href="https://www.weamec.fr/en/projects/helofow/">the WEAMEC website</a>. </p>
Dataset for the paper "Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset"
<p>We present a large-scale anomaly detection dataset collected from IBM Cloud's Console over approximately 4.5 months. This high-dimensional dataset captures telemetry data from multiple data centers, specifically designed to aid researchers in developing and benchmarking anomaly detection methods in large-scale cloud environments. It contains 39,365 entries, each representing a 5-minute interval, with 117,448 features/attributes, as interval_start is used as the index. The dataset includes detailed information on request counts, HTTP response codes, and various aggregated statistics. The dataset also includes labeled anomaly events identified through IBM's internal monitoring tools, providing a comprehensive resource for real-world anomaly detection research and evaluation.</p> <p><strong>File Descriptions</strong></p> <ul> <li><code>location_downtime.csv</code> - Details planned and unplanned downtimes for IBM Cloud data centers, including start and end times in ISO 8601 format.</li> <li><code>unpivoted_data.parquet</code> - Contains raw telemetry data with 413 million+ rows, covering details like location, HTTP status codes, request types, and aggregated statistics (min, max, median response times).</li> <li><code>anomaly_windows.csv</code> - Ground truth for anomalies, listing start and end times of recorded anomalies, categorized by source (Issue Tracker, Instant Messenger, Test Log).</li> <li><code>pivoted_data_all.parquet</code> - Pivoted version of the telemetry dataset with 39,365 rows and 117,449 columns, including aggregated statistics across multiple metrics and intervals.</li> <li><code>demo/demo.[ipynb|html]</code>: This demo file provides examples of how to access data in the Parquet files, available in Jupyter Notebook (<code>.ipynb</code>) and HTML (<code>.html</code>) formats, respectively.</li> </ul> <p>Further details of the dataset can be found in <strong>Appendix B: Dataset Characteristics</strong> of the <a href="https://arxiv.org/abs/2411.09047">paper</a> titled <strong><em>"Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset."</em></strong> Sample code for training anomaly detectors using this data is provided in <a href="https://doi.org/10.5281/zenodo.14598119" target="_blank" rel="noopener">this package</a>.</p> <p> </p> <p>When using the dataset, please cite it as follows:</p> <pre><code>@misc{islam2024anomaly,</code><br><code> title={Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset}, </code><br><code> author={Mohammad Saiful Islam and Mohamed Sami Rakha and William Pourmajidi and Janakan Sivaloganathan and John Steinbacher and Andriy Miranskyy},</code><br><code> year={2024},</code><br><code> eprint={2411.09047},</code><br><code> archivePrefix={arXiv},</code><br><code> url={https://arxiv.org/abs/2411.09047}</code><br><code>}</code></pre> <p> </p>
One-hectare fine-scale dataset of a fynbos plant community in the Cape Floristic Region
<p>Cape fynbos, which forms part of the Cape Floristic Region (CFR) of South Africa, a global biodiversity hotspot, is renowned for its high levels of plant species endemism and diversity. This extraordinary ecosystem, characterised by nutrient-poor soils and fire-adapted vegetation, is a treasure trove of endemic flora. However, this fragile system faces increasing threats from habitat loss, climate change, and invasive species. Pristine fynbos, naturally high in plant diversity and which forms a large part of the CFR, presents an ideal opportunity to gather fine-scale data on community assembly patterns. Most fynbos vegetation surveys use a plot size of about 100 m2, with no spatial structures within plots to demarcate individual subplots. Here, a groundbreaking dataset is presented that fully covers 1-hectare of pristine fynbos, systematically gridded into 50 × 50 subplots, each measuring 2 × 2 m, arranged evenly within a square-shaped survey site. Each plot was assigned a unique Y–X coordinate combination. For each plot, all plant species present were recorded, along with their total percentage covers and maximum height values. Total percentage covers were also recorded for bare soil, rock, and termite mounds. This dataset provides a valuable contribution to the field of fynbos ecology, as well as plant community ecology in general, and establishes a benchmark for future one-hectare surveys of similar fynbos vegetation types, delineating the fine-scale composition and structure of fynbos in the CFR. The dataset will be useful for a wide audience, including community and spatial ecologists, plant and environmental scientists, and biodiversity informaticians and statistical ecologists, offering ideal data for testing new metrics of diversity and compositional turnover. Data in Brief, Volume 59, April 2025, 111334: <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.dib.2025.111334" target="_blank" rel="noreferrer noopener"><span><span>https://doi.org/10.1016/j.dib.2025.111334</span></span></a></p>
Chemical sensors for fire detection and nuisance rejection under EN-5420 standard conditions and reduced-scale chamber
<p>The dataset was acquired using a gas sensor array placed in the celling of a validated standard fire room (240 m3) located in Minimax Company. The dataset includes measurements of three different campaigns that were performed over 15 months. The dataset includes standard EN-54 smoldering fires and non-standard smoldering fires (such as plastic fires; PVC, cables Fire). In order to generate scenarios that may result in false-positive alarms when gas sensors are used, different nuisance experiments were also performed (such as cleaners, and air fresheners). Additionally, an additional measurement campaign was performed in a small chamber. The small-scale experiments dataset includes scale-down replicates of the fire and nuisances experiments performed in the standard fire room (EN-54 smoldering fire experiments, non-standard fires, and nuisance experiments).</p> <p>Citation request: Ana Solórzano et al, Early fire detection based on gas sensor arrays: Multivariate calibration and validation, Sensors and Actuators B: Chemical, 2021, <a href="https://doi.org/10.1016/j.snb.2021.130961">https://doi.org/10.1016/j.snb.2021.130961</a>.</p>
Clonal decomposition and DNA replication states defined by scaled single cell genome sequencing
<p><strong>OV2295 Tables</strong></p> <p>ov2295_breakpoint_counts.csv.gz: Table of breakpoint counts per cell</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>cell_id: identifier for the cell</li> <li>read_count: number of reads</li> <li>library_id: identifier for the DNA library</li> <li>sample_id: identifier for the sequenced sample</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> </ul> <p>ov2295_cell_cn.csv.gz: Table of cell specific copy number</p> <ul> <li>cell_id: identifier for the cell</li> <li>sample_id: identifier for the sequenced sample</li> <li>library_id: identifier for the DNA library</li> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>reads: number of reads</li> <li>copy: raw normalized copy number</li> <li>state: copy number state</li> <li>gc: percent gc of the bin</li> <li>map: average mappability of the bin</li> </ul> <p>ov2295_cell_metrics.csv.gz: Table of cell metrics</p> <ul> <li>cell_id: identifier of the cell</li> <li>unpaired_mapped_reads: number of unpaired mapped reads</li> <li>paired_mapped_reads: number of mapped reads that were properly paired</li> <li>unpaired_duplicate_reads: number of unpaired duplicated reads</li> <li>paired_duplicate_reads: number of paired reads that were also marked as duplicate</li> <li>unmapped_reads: number of unmapped reads</li> <li>percent_duplicate_reads: percentage of duplicate reads</li> <li>estimated_library_size: scaled total number of mapped reads</li> <li>total_reads: total number of reads, regardless of mapping status</li> <li>total_mapped_reads: total number of mapped reads</li> <li>total_duplicate_reads: number of duplicate reads</li> <li>total_properly_paired: number of properly paired reads</li> <li>coverage_breadth: percentage of genome covered by some read</li> <li>coverage_depth: average reads per nucleotide position in the genome</li> <li>median_insert_size: median insert size between paired reads</li> <li>mean_insert_size: mean insert size between paired reads</li> <li>standard_deviation_insert_size: standard deviation of the insert size between paired reads</li> <li>index_sequence: index sequence of the adaptor sequence</li> <li>column: column of the cell on the nanowell chip</li> <li>img_col: column of the cell from the perspective of the microscope</li> <li>index_i5: id of the i5 index adapter sequence</li> <li>sample_type: type of the sample</li> <li>primer_i7: id of the i5 index primer sequence</li> <li>experimental_condition: experimental treatment of the cell, includes controls</li> <li>index_i7: id of the i7 index adapter sequence</li> <li>cell_call: living/dead classification of the cell based on staining usually, C1 == living, C2 == dead</li> <li>sample_id: name of the sample</li> <li>primer_i5: id of the i5 index primer sequence</li> <li>row: row of the cell on the nanowell chip</li> <li>library_id: identifier for the DNA library</li> <li>index: ignored</li> <li>multiplier: during parameter searching, the set [1..6] that was chosen</li> <li>MSRSI_non_integerness: median of segment residuals from segment integer copy number states</li> <li>MBRSI_dispersion_non_integerness: median of bin residuals from segment integer copy number states</li> <li>MBRSM_dispersion: median of bin residuals from segment median copy number values</li> <li>autocorrelation_hmmcopy: hmmcopy copy autocorrelation</li> <li>cv_hmmcopy: ignored</li> <li>empty_bins_hmmcopy: number of empty bins in hmmcopy</li> <li>mad_hmmcopy: median absolute deviation of hmmcopy copy</li> <li>mean_hmmcopy_reads_per_bin: mean reads per hmmcopy bin</li> <li>median_hmmcopy_reads_per_bin: median reads per hmmcopy bin</li> <li>std_hmmcopy_reads_per_bin: standard deviation value of reads in hmmcopy bins</li> <li>total_halfiness: summed halfiness penality score of the cell</li> <li>total_mapped_reads_hmmcopy: total mapped reads in all hmmcopy bins</li> <li>scaled_halfiness: summed scaled halfiness penalty score of the cell</li> <li>mean_state_mads: mean value for all median absolute deviation scores for each state</li> <li>mean_state_vars: variance value for all median absolute deviation scores for each state</li> <li>mad_neutral_state: median absolute deviation score of the neutral 2 copy state</li> <li>breakpoints: number of breakpoints, as indicated by state changes not at the ends of chromosomes</li> <li>mean_copy: mean hmmcopy copy value</li> <li>state_mode: the most commonly occuring state</li> <li>log_likelihood: hmmcopy log likelihood for the cell</li> <li>true_multiplier: the exact decimal value used to scale the copy number for segmentation</li> <li>order: order of the cell in the hierarchical clustering tree</li> <li>quality: random forest classifier proability score that cell is good</li> </ul> <p>ov2295_clone_alleles.csv.gz: Table of clone specific allele data</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>hap_label: haplotype block identifier</li> <li>clone_id: clone identifier</li> <li>allele_1_sum: number of reads for allele 1 of the haplotype block</li> <li>allele_2_sum: number of reads for allele 2 of the haplotype block</li> <li>total_counts_sum: total reads for the haplotype block</li> </ul> <p>ov2295_clone_breakpoints.csv.gz: Table of breakpoints per clone for OV2295 samples. Columns:</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> <li>clone_id: clone identifier</li> <li>read_count: number of reads</li> <li>is_present: presence=1, absent=0</li> </ul> <p>ov2295_clone_clusters.csv.gz: Table of cell clusters as putative clones</p> <ul> <li>cell_id: identifier for the cell</li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_cn.csv.gz: Table of allele specific copy number per clone for OV2295 samples. Columns:</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>total_cn: HMMCopy predicted total copy number </li> <li>minor_cn: HMM predicted minor copy number </li> <li>major_cn: HMM predicted major copy number </li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_snvs.csv.gz: Table of SNVs per clone for OV2295 samples. Columns:</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>clone_id: clone identifier</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>total_counts: total number of reads at this position</li> <li>is_present: presence=0, absent=1</li> <li>is_het: is heterozygous</li> <li>is_hom: is homozygous for the alternate</li> </ul> <p>ov2295_nodes.csv.gz: Table of phylogenetic information for SNV evolution</p> <ul> <li>variant_id: identifier for the SNV as chrom:coord:ref:alt</li> <li>node: node in the phylogenetic tree</li> <li>loss: probability the SNV was lost at this node</li> <li>origin: probability the SNV originated at this node</li> <li>presence: probability the SNV is present at this node</li> <li>ml_origin: binary indicator the SNV originated at this node</li> <li>ml_presence: binary indicator the SNV is present at this node</li> <li>ml_loss: binary indicator the SNV was lost at this node</li> </ul> <p>ov2295_snv_counts.csv.gz: Table of SNV counts</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>cell_id: identifier for the cell</li> <li>total_counts: total number of reads at this position</li> <li>sample_id: identifier for the sequenced sample</li> </ul> <p>ov2295_tree.pickle: Phylogenetic tree in python pickle format. Requires installation of the stochastic dollo code at: https://bitbucket.org/dranew/dollo, version 0.4.2.</p> <p>Note the following sample mapping: ‘SA922’: ‘OV2295(R2)’, ‘SA921’: ‘TOV2295(R)’, ‘SA1090’: ‘OV2295’,</p> <p><strong>Plots</strong></p> <p>ov_supp_clone_allele_cn.png: Clone allele ratios for each OV2295 sample.</p> <p>ov_supp_clone_total_cn.png: Clone copy number for each OV2295 sample.</p> <p>ov_supp_sample_total_cn.png: Bulk copy number for each OV2295 sample.</p> <p>ov_supp_sample_allele_cn.png: Bulk allele ratios for each OV2295 sample.</p>
Massive IoT for Large-Scale Public Events in the 5GENESIS Surrey Platform
<p>This dataset contains the results of the trials conducted within the context of the main IoT use case of the 5GENESIS Surrey Platform.</p>
Dataset to manuscript: Soil organic carbon stocks and quality in small-scale tropical, sub-humid and semi-arid watersheds under shrubland and dry deciduous forest in southwestern India
<p>Raw data to the manuscript entitled "Soil organic carbon stocks and quality in small-scale tropical, sub-humid and semi-arid watersheds under shrubland and dry deciduous forest in southwestern India" by Severin-Luca Bellè, Jean Riotte, Muddu Sekhar, Laurent Ruiz, Marcus Schiedung and Samuel Abiven.</p> <p>Data files include all raw data of soil cores (20211111_Raw_data.zip), data measured on composited samples (20211111_Composite_data.zip) and DRIFT spectra (20211111_DRIFT_data.zip).</p> <p>Files ending with var_names are the README files.</p>
Supplementary data: Accurate large-scale simulations of siliceous zeolites by neural network potentials
<p><strong>Content</strong></p> <p><em>1. Zeolite databases</em></p> <ul> <li>Deem database containing 331170 hypothetical zeolite frameworks [Deem09, Pophale11] geometrically optimized at the NNPscan level (note, the first row of the database is alpha-quartz): "DEEM_NNPscan.db"</li> <li>Database of 236 exiting zeolite frameworks of the <a href="http://www.iza-structure.org/databases/">International Zeolite Association (IZA) </a>optimized at the NNPscan level: "IZA_NNPscan.db"</li> <li>Both databases are <a href="https://wiki.fysik.dtu.dk/ase/ase/db/db.html">ASE SQLite database files</a> of the <a href="https://wiki.fysik.dtu.dk/ase/index.html">Atomic Simulation Environment</a> containing the ASE <a href="https://wiki.fysik.dtu.dk/ase/ase/atoms.html">Atoms objects</a> with energies and forces (NNPscan level); readable with ASE's <a href="https://wiki.fysik.dtu.dk/ase/ase/io/io.html">I/O module</a></li> <li>Additionally, relevant quantities can be extracted with, e.g., the following queries (further information: ase db --help):</li> </ul> <pre><code class="language-bash">ase db DEEM_NNPscan.db -c id,formula,natoms,volume,mass,density,energy_per_tsite,n_tsites,relative_energy # Output id|formula|natoms| volume| mass|density|energy_per_tsite|n_tsites|relative_energy 1|O6Si3 | 9|111.161|180.249| 26.988| -31.796| 3| 0.000 2|O16Si8 | 24|433.858|480.664| 18.439| -31.638| 8| 15.265 3|O16Si8 | 24|421.114|480.664| 18.997| -31.596| 8| 19.359 4|O16Si8 | 24|426.557|480.664| 18.755| -31.614| 8| 17.613 5|O16Si8 | 24|412.410|480.664| 19.398| -31.613| 8| 17.677 6|O16Si8 | 24|393.544|480.664| 20.328| -31.594| 8| 19.546 7|O16Si8 | 24|422.400|480.664| 18.939| -31.657| 8| 13.476 8|O16Si8 | 24|394.405|480.664| 20.284| -31.581| 8| 20.797 9|O12Si6 | 18|265.201|360.498| 22.624| -31.611| 6| 17.868 10|O16Si8 | 24|357.047|480.664| 22.406| -31.581| 8| 20.785 11|O16Si8 | 24|434.894|480.664| 18.395| -31.621| 8| 16.911 12|O16Si8 | 24|384.158|480.664| 20.825| -31.657| 8| 13.448 13|O12Si6 | 18|258.977|360.498| 23.168| -31.679| 6| 11.278 14|O16Si8 | 24|466.429|480.664| 17.152| -31.593| 8| 19.588 15|O16Si8 | 24|423.469|480.664| 18.892| -31.639| 8| 15.179 16|O16Si8 | 24|450.716|480.664| 17.750| -31.628| 8| 16.219 17|O16Si8 | 24|331.528|480.664| 24.131| -31.642| 8| 14.857 18|O16Si8 | 24|458.573|480.664| 17.445| -31.635| 8| 15.572 19|O16Si8 | 24|359.298|480.664| 22.266| -31.655| 8| 13.636 20|O16Si8 | 24|464.264|480.664| 17.232| -31.612| 8| 17.750 Rows: 331171 (showing first 20) Keys: density, energy_per_tsite, n_tsites, relative_energy ase db IZA_NNPscan.db -c id,formula,natoms,volume,mass,density,energy_per_tsite,n_tsites,relative_energy,iza_code # Output id|formula |natoms| volume| mass|density|energy_per_tsite|n_tsites|relative_energy|iza_code 1|O16Si8 | 24| 435.488| 480.664| 18.370| -31.676| 8| 11.594|ABW 2|O32Si16 | 48| 961.419| 961.328| 16.642| -31.645| 16| 14.612|ACO 3|O96Si48 | 144|3154.579|2883.984| 15.216| -31.664| 48| 12.810|AEI 4|O80Si40 | 120|2102.921|2403.320| 19.021| -31.703| 40| 9.021|AEL 5|O96Si48 | 144|2417.286|2883.984| 19.857| -31.666| 48| 12.586|AEN 6|O144Si72| 216|4075.300|4325.976| 17.667| -31.674| 72| 11.831|AET 7|O96Si48 | 144|2786.810|2883.984| 17.224| -31.675| 48| 11.716|AFG 8|O48Si24 | 72|1400.247|1441.992| 17.140| -31.690| 24| 10.268|AFI 9|O64Si32 | 96|1764.823|1922.656| 18.132| -31.653| 32| 13.809|AFN 10|O80Si40 | 120|2080.330|2403.320| 19.228| -31.707| 40| 8.632|AFO 11|O64Si32 | 96|2097.384|1922.656| 15.257| -31.655| 32| 13.622|AFR 12|O112Si56| 168|3820.116|3364.648| 14.659| -31.650| 56| 14.150|AFS 13|O144Si72| 216|4732.720|4325.976| 15.213| -31.664| 72| 12.793|AFT 14|O60Si30 | 90|1897.074|1802.490| 15.814| -31.659| 30| 13.268|AFV 15|O96Si48 | 144|3154.885|2883.984| 15.214| -31.664| 48| 12.776|AFX 16|O32Si16 | 48|1137.335| 961.328| 14.068| -31.591| 16| 19.790|AFY 17|O48Si24 | 72|1283.812|1441.992| 18.694| -31.620| 24| 17.034|AHT 18|O96Si48 | 144|2479.287|2883.984| 19.360| -31.681| 48| 11.155|ANA 19|O64Si32 | 96|1797.086|1922.656| 17.807| -31.662| 32| 12.924|APC 20|O64Si32 | 96|1751.393|1922.656| 18.271| -31.678| 32| 11.422|APD Rows: 236 (showing first 20) Keys: density, energy_per_tsite, iza_code, n_tsites, relative_energy # Filtering of the database, e.g., for structures with relative energies < 10 kJ/(mol Si) ase db IZA_NNPscan.db relative_energy\<10 -c density,energy_per_tsite,n_tsites,relative_energy,iza_code # Output density|energy_per_tsite|n_tsites|relative_energy|iza_code 19.021| -31.703| 40| 9.021|AEL 19.228| -31.707| 40| 8.632|AFO 19.385| -31.695| 24| 9.802|ATV 18.778| -31.702| 34| 9.061|DOH 19.570| -31.693| 24| 9.959|EWO 18.401| -31.698| 32| 9.451|GON 18.551| -31.695| 112| 9.807|IHW 17.778| -31.693| 288| 9.972|IMF 19.154| -31.695| 6| 9.762|JBW 18.187| -31.695| 96| 9.734|MFI 19.278| -31.709| 48| 8.443|MRE 18.035| -31.698| 90| 9.481|MSO 20.417| -31.724| 44| 7.003|MTF 19.227| -31.704| 136| 8.898|MTN 18.542| -31.693| 28| 9.966|MTW 19.137| -31.695| 60| 9.798|PCR 20.037| -31.709| 144| 8.464|PSI 18.843| -31.703| 64| 9.004|SAF 18.371| -31.703| 112| 8.975|STO 19.894| -31.706| 17| 8.671|VET Rows: 20 (showing first 20) Keys: density, energy_per_tsite, iza_code, n_tsites, relative_energy</code></pre> <ul> <li>The quantities shown above are available with the keys (besides standard ASE database keys):</li> </ul> <table> <thead> <tr> <th scope="col">Key</th> <th scope="col">Quantity</th> <th scope="col">Unit</th> </tr> </thead> <tbody> <tr> <td>id</td> <td>Identifier</td> <td> </td> </tr> <tr> <td>formula</td> <td>Chemical formula of the unit cell</td> <td> </td> </tr> <tr> <td>natoms</td> <td>Number of atoms</td> <td> </td> </tr> <tr> <td>volume</td> <td>Unti cell volume</td> <td>Å<sup>3</sup></td> </tr> <tr> <td>mass</td> <td>Atomic mass of the unit cell</td> <td>amu</td> </tr> <tr> <td>density</td> <td>Framework density</td> <td>Si/nm<sup>3</sup></td> </tr> <tr> <td>energy_per_tsite</td> <td>NNPscan energy</td> <td>eV</td> </tr> <tr> <td>n_tsites</td> <td>Number of T-sites</td> <td> </td> </tr> <tr> <td>relative_energy</td> <td>Energy with respect to quartz</td> <td>kJ/(mol Si)</td> </tr> <tr> <td>iza_code</td> <td>only for 'IZA_NNPscan.db'</td> <td> </td> </tr> </tbody> </table> <ul> <li> Comma separated csv files for the quantities listed above: "DEEM_NNPscan.csv" and "IZA_NNPscan.csv"</li> </ul> <p><em>2. Neural network potentials (NNP) for silica</em></p> <ul> <li>SchNet [Schütt18,Schütt19] NNP files trained on DFT data at the PBE+D3 (NNPpbe) and SCAN+D3 level (NNPscan)</li> <li>Simulations can be performed using <a href="https://schnetpack.readthedocs.io/en/stable/getstarted/getstarted.html#references">SchNetPack</a> with its ASE calculator</li> <li>This example shows a simple single-point calculation</li> </ul> <pre><code class="language-python">import ase.io import torch from schnetpack.interfaces import SpkCalculator from schnetpack.environment import AseEnvironmentProvider # check if GPU(s) are available if torch.cuda.is_available(): device = "cuda" else: device = "cpu" # load the NNP model model = torch.load('SiOscan1', map_location=device) # read some structure atoms = ase.io.read( ... ) # define SchNetPack calculator calc = SpkCalculator(model=model, device=device, energy='energy', forces='forces', environment_provider=AseEnvironmentProvider(6.) ) # attach calculator to atoms object atoms.set_calculator(calc) # perform simulations, e.g., single-point calculation energy = atoms.get_potential_energy() print(energy)</code></pre> <p><em>3. Test set used for accuracy evaluation (ASE database: test_set_NNPscan.db)</em></p>
Experimental data for 'Scaling laws for coastal overwash morphology'
<p>This dataset contains the experimental data described in Lazarus, ED (2016) Scaling laws for coastal overwash morphology, <em>Geophysical Research Letters</em>, 43, 12113–12119, <a href="https://doi.org/10.1002/2016GL071213">https://doi.org/10.1002/2016GL071213</a>.</p> <p>The physical experiments that produced these data were conducted at St Anthony Falls Laboratory (University of Minnesota, USA) in December 2014. The experiments were conducted in a 3 x 5 x 0.6 m tank filled with well-sorted coarse river sand. The tank and the experimental trials are detailed in Text S1 of the Supporting Information for Lazarus (2016): <a href="https://agupubs.onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1002%2F2016GL071213&file=grl55284-sup-0001-SI.pdf">https://agupubs.onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1002%2F2016GL071213&file=grl55284-sup-0001-SI.pdf</a></p> <p>This dataset consists of two *.csv files:</p> <ul> <li>'...THROATS.csv' – morphometric data for <strong>erosional</strong> (throat) features in the experimental barrier</li> <li>'...WASHOVER.csv' – morphometric data for <strong>depositional</strong> (washover) features on the back-barrier floodplain</li> </ul> <p>Both files have the same general column headings: feature width (in the alongshore dimension) [m], feature length (in the cross-shore dimension) [m], feature area [m<sup>2</sup>], feature volume [m<sup>3</sup>], alongshore spacing (centroid-to-centroid distance to neighbouring feature) [m], and real alongshore position [m].</p> <p>All features were formed along an initially geometrically uniform (topographically homogenous) trapezoidal barrier under inundation-type forcing (denoted in 'forcing' column). These data report the compiled results of three experimental trials (denoted in 'trial' column).</p> <p>Note that these data are also available as part of the Supporting Information for Lazarus (2016), but the format in which they were originally uploaded is not conducive to straightforward integration into open-source analysis. Publishing them here, in this tidier format, is an effort to rectify that.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.