Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,326
datasets available to search
ShareScore release 0.9.0
Dataset results
2,326 results for “Clustering”
Equation-of-Motion Coupled-Cluster Theory based on the 4-component Dirac–Coulomb(–Gaunt) Hamiltonian. Energies for single electron detachment, attachment and electronically excited states: Figures
<p>This entry contains the figures included in the paper titled "Equation-of-Motion Coupled-Cluster Theory based on the 4-component Dirac--Coulomb(--Gaunt) Hamiltonian. Energies for single electron detachment, attachment and electronically excited states", by Avijit Shee, Trond Saue, Lucas Visscher and Andre Severo Pereira Gomes.</p> <p>It accompanies the dataset found at the DOI: 10.5281/zenodo.1320320</p> <p>There are three figures that use the (original) png files included in <a href="https://zenodo.org/api/files/7bda2e2b-ac69-41aa-a21e-821e88bfb973/original-figures.tar.bz2">original-figures.tar.bz2 </a>:</p> <p>figure 1: Potential energy curves of the spin-orbit split X<sup>2</sup>Π and A<sup>2</sup>Π states of the XO molecules, obtained with EOM-IP and the <sup>2</sup>DCG<sup>M</sup> Hamiltonian.</p> <p>figure 2: Internuclear distances (in Angstrom), harmonic vibrational frequencies (in cm<sup>−1</sup>) and the vertical Ω = 3/2 − 1/2 energy difference (in eV) for the X<sup>2</sup>Π and A<sup>2</sup>Π states of the XO molecules, obtained with EOM-IP and the <sup>2</sup>DCG<sup>M</sup> Hamiltonian.</p> <p>figure 3: SO-ZORA/QZ4P/Hartree-Fock (ADF) spinor magnetization plots (isosurfaces at 0.03 a.u.) and energies (in Eh) for the valence spinors of the XO<sup>−</sup> species (from left to right: X = Cl, Br, I, At, Ts).</p>
Predictive simulations of ionization energies of solvated halide ions with relativistic embedded Equation of Motion Coupled-Cluster Theory: Figures
<p>This entry contains the sources for the figures included in the body of the paper titled "Predictive simulations of ionization energies of solvated halide ions with relativistic embedded Equation of Motion Coupled-Cluster Theory", by Yassine Bouchafra, Avijit Shee, Florent Réal, Valérie Vallet and André Severo Pereira Gomes, as well as those found in the supplementary information.</p> <p>It accompanies the dataset found at the DOI: 10.5281/zenodo.1477004</p> <p> </p> <p> </p>
Supplementary material for the paper "DS Andromedae, A Detached Eclipsing Double-Lined Spectroscopic Binary in the Galactic Cluster NGC 752
<p>Supplementary material supporting the paper "DS Andromedae: A Detached Eclipsing Double-Lined Spectroscopic</p> <p>Binary in the Galactic Cluster NGC 752" by E. F. Milone, S. J. Schiller, Th. Mellergaard Amby, and S. Frandsen.</p> <p>It includes:</p> <p>A Read-me file in three formats (docx, rtf, pdf); Unabridged Section 3 with extended modeling details (pdf);</p> <p>Extended spreadsheet version of Table 3 of adjusted parameters (pdf); Extended spreadsheet version of Table 8 of absolute</p> <p>parameters (pdf); and Complete Table 15 of photometric data (txt): and a sample DC input file (for Model 41, used in the</p> <p>DS And modeling) in dat format.</p>
Data: Dynamics of star clusters with tangentially anisotropic velocity distribution (Pavlik+ 2024)
<p>This dataset represents the results of our <em>N</em>-body simulations of star clusters (SCs). The initial conditions of the models are fully described in the referenced journal article. In short, the SCs start from isotropic, radially anisotropic or tangentially anisotropic initial velocity distributions, and each model is evolved in an external Galactic tidal field, for two different choices of the filling factor.</p>
Energy recovery by an unbiased gas phase photofuel cell with a nickel foam supported WO3 photoanode decorated with plasmonic gold clusters
<p>Dataset for the article titled "Energy recovery by an unbiased gas phase photofuel cell with a nickel foam supported WO3 photoanode decorated with plasmonic gold clusters".</p> <p>This research was conducted at Antwerp Engineering, Photoelectrochemistry and Sensing (A-PECS) group, University of Antwerp, Belgium.</p>
GFN2-xTB structures of iCOM adsorbed on a cluster model of water molecules derived from a periodic model of crystalline ice
<p>This dataset contains the atomic coordinates in the <a href="http://www.moldraw.unito.it/">.</a>xyz format of the GFN2-xTB optimized structures of 20 iCOMs adsorbed at the surface of a cluster of 84 water molecules mimicking the periodic model of crystalline water icy grain as described by Ferrero, S.; Zamirri, L., Ceccarelli, C.; Witzel, A.; Rimola, A.; Ugliengo, P. ApJ, (2020) 904:11. For all considered structures we also provided a specific file in the Gaussian format with the computed harmonic frequencies. Each file can be easily converted in input for the variety of quantum mechanical programs, like VASP, QE, Gaussian 16 etc.</p> <p> </p> <p> </p>
Offering ART refill through community health workers versus clinic-based follow-up after home-based same-day ART initiation in rural Lesotho: The VIBRA cluster-randomised clinical trial
<p>These are pseudo-anonymised data from the VIBRA randomized trial: "Offering ART refill through community health workers versus clinic-based follow-up after home-based same-day ART initiation in rural Lesotho: The VIBRA cluster-randomised clinical trial". The data dictionary explains the data available in the dataset. Between August 2018 and May 2019, 257 eligible individuals from 117 consenting villages were enrolled from two districts of Lesotho, and followed up for a maximum of 15 months. The protocol was published, doi: 10.1186/s13063-019-3510-5</p>
Supporting data for manuscript "Geochemical Characterization of Insoluble Particle Clusters in Ice Cores Using Two-dimensional Impurity Imaging"
<p>Laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS) offers micron-resolution 2D chemical imaging, which has been adapted recently to ice core analysis. The datasets are supporting information for the manuscript "Geochemical Characterization of Insoluble Particle Clusters in Ice Cores Using Two-dimensional Impurity Imaging" accepted for publication at Geochemistry, Geophysics, Geosystems (10.1029/2022GC010595). Measurements were performed at the Ca’Foscari University of Venice, considered as analytes are 23Na, 24Mg, 27Al, 29Si, 43Ca, 56Fe and 88Sr. Background and drift correction as well as image construction were performed using the software HDIP (Teledyne Photon Machines, Bozeman, MT, USA). Impurity maps are acquired as a pattern of lines, without overlap in the direction perpendicular to that of the scan, and without any further spatial interpolation. In a sample of the EGRIP Greenland ice core (from about 1256.95 m depth), maps were obtained over 3 adjacent areas. For each of the maps, for every chemical channel the intensities (in counts, after background and drift correction) are provided as a separate file, named as “ds01_Area1_Na.csv”, etc. These maps were obtained using a 20 µm square spot. This data can be used to obtain the images shown in the manuscript. For the additional map shown as Figure 9 in the manuscript, data were obtained using a LA-ICP-TOFMS for imaging a sample of the last glacial period in the EPICA Dome C (EDC) ice core, bag 1065. The maps were acquired using a 35 µm square spot, with 50% overlap between neighboring pixels to increase the spatial resolution horizontally.</p>
CLUSTER anti-TNF clinical dataset
<p>This data dictionary describes the CLUSTER anti-TNF clinical dataset. It contains key variables of interest for measuring treatment response to tumour necrosis factor inhibitors (anti-TNF) in juvenile idiopathic arthritis (JIA) patients at baseline/timepoint 1 (start of anti-TNF treatment or up to 3 months before) and timepoint 2 (6 months after baseline, within a 3-12 month window). If a patient had multiple visits within a timepoint window, the one closest to baseline/6 months was selected. There are 2401 patients in this dataset.</p> <p>CLUSTER has brought together data from 4 UK observational JIA research cohort studies: the UK JIA Biologics Register, the Childhood Arthritis Prospective Study (CAPS) and the Childhood Arthritis Response to Medication Study (CHARMS). Clinical data from these studies was pooled to create the CLUSTER anti-TNF clinical dataset. More details of CLUSTER's clinical data harmonisation process are available here: <link></p> <p>We are open to collaborating and sharing data with researchers who have research questions that may be answered by our data. Access to CLUSTER data is regulated according to the conditions of patient consent, study ethics and the CLUSTER Consortium’s policies. Access is available to researchers via an application to the CLUSTER data access committee - more details are available on <a href="https://www.clusterconsortium.org.uk/">CLUSTER's website.</a></p>
CLUSTER MTX Clinical Dataset
<p>This data dictionary describes the CLUSTER MTX clinical dataset. It contains key variables of interest for measuring treatment response to methotrexate (MTX) in juvenile idiopathic arthritis (JIA) patients at baseline/timepoint 1 (start of MTX treatment or up to 3 months before) and timepoint 2 (6 months after baseline, within a 3-12 month window). If a patient had multiple visits within a timepoint window, the one closest to baseline/6 months was selected. There are 2899 patients in this dataset.</p> <p>CLUSTER has brought together data from 4 UK observational JIA research cohort studies: the UK JIA Biologics Register, the Childhood Arthritis Prospective Study (CAPS) and the Childhood Arthritis Response to Medication Study (CHARMS). Clinical data from these studies was pooled to create the CLUSTER MTX clinical dataset. More details of CLUSTER's clinical data harmonisation process are available here: <link></p> <p>We are open to collaborating and sharing data with researchers who have research questions that may be answered by our data. Access to CLUSTER data is regulated according to the conditions of patient consent, study ethics and the CLUSTER Consortium’s policies. Access is available to researchers via an application to the CLUSTER data access committee - more details are available on <a href="https://www.clusterconsortium.org.uk">CLUSTER's website.</a></p>
Dataset from the paper "Eccentric black hole mergers via three-body interactions in young, globular and nuclear star clusters"
<p>This repository contains several data from the paper "Eccentric black hole mergers via three-body interactions in young, globular and nuclear star clusters".</p> <p> </p> <p><strong>BBH_mergers_cat_*.dat</strong> contains the data for the BBH merger population produced by the three-body simulations. These data can be used to reproduce figures 4,5, and 8 of the paper. The file is organized in columns as:</p> <ul> <li>ID of the simulation.</li> <li>outcome of the simulation (12, merger triggered by a flyby event, 13 and 23 merger triggered after an exchange event in which the secondary (primary) BH is replaced by the intruder, 123 second generation BBH merger.</li> <li>mass of the primary BH in solar masses</li> <li>mass of the secondary BH in solar masses</li> <li>Chirp mass of the system in solar masses</li> <li>coalescence time since the beginning of the simulation in year (note that all the simulation with tcoal<1e5 yr have merged during the direct N-body simulation, while all the mergers that take place after this value are evolved with the equations by Peters 1964)</li> <li>eccentricity of the binary at 10 Hz in the detector frame</li> <li>tilt angle in radiant, defined as the angle between the orbital plane of the initial binary at the beginning of the simulation and the orbital plane of the final binary at the end of the simulation.</li> </ul> <p>The files named <strong>data_*.txt</strong> contains the masses, the position and the velocities at each timestep for the three simulations showed in fig.1 in the paper. The data are referred to the center-of-mass of the three-body system. The file is organized as follows:</p> <ul> <li>The first line of the file reports the masses in solar masses of the three BHs.</li> <li>Column 0 reports the time in yr</li> <li>Colum 1-3 report the x,y,z position for the m1 BH in parsec</li> <li>Colum 4-6 report the x,y,z position for the m2 BH in parsec</li> <li>Colum 7-9 report the x,y,z position for the m3 BH in parsec</li> <li>Colum 10-12 report the x,y,z components of the velocities of the m1 BH in km/s</li> <li>Colum 13-15 report the x,y,z components of the velocities of the m1 BH in km/s</li> <li>Colum 16-18 report the x,y,z components of the velocities of the m1 BH in km/s</li> </ul> <p>Finally, <strong>outcomes_*.dat</strong> contains two columns:</p> <ul> <li>Column 0 reports the ID of the simulation</li> <li>Column 1 reports the outcome of the simulation as: 12 flyby (or merger after a flyby), 13 and 23 exchange (or merger after an exchange) in which the secondary (primary) BH is replaced by the intruder, 0 in the system is ionized in three single BHs, 3 if the system is still interacting at 1Myr, i.e. when we stop our simulation.</li> </ul> <p>This file might be useful to train a machine-lerning classificator, and can be used to reproduce Fig. 2 of the paper.</p> <p> </p> <p><strong>Contacts:</strong></p> <p>Marco Dall'Amico</p> <p>marco.dallamico@phd.studenti.unipd.it</p> <p>marco.dallamico@pd.infn.it</p>
Application of two-step clustering algorithm to QuaLiKiz-v2.6.2 turbulent transport simulation data
<p>QuaLiKiz simulation data in support of the two-step clustering algorithm, developed by Bart J. J. Kremers.</p> <p>The NETCDF file, generated via NETCDF4, contains the raw QuaLiKiz output for the 3-dimensional (2-input, 1-output) toy case used to develop the algorithm. Within the NETCDF file, the coordinates represent the code inputs and various vector indices and the data variables represent the code outputs.</p> <p>There are also 4 HDF5 files, containing the results from the two-step clustering reduction algorithm, where the data is saved under 2 keys: "/input" and "/flattened". The file names indicate the reduction algorithm settings used to produce the results within.</p> <p>The algorithm is available open-source at <a href="https://gitlab.com/BartKremers/two-step-clustering">https://gitlab.com/BartKremers/two-step-clustering</a>.</p>
Integrated freshwater abundance and connectivity clusters at the Hydrologic Unit 8 scale for the Midwest and Northeast U.S.A. – freshwater metric variables and k-means cluster assignment
This dataset includes integrated freshwater abundance and connectivity cluster output, principal component scores, and lake, wetland, and stream abundance and connectivity metrics measured at the Hydrologic Unit 8 (HU8) scale for 17 U.S. states in the Midwest and Northeast regions (appr. 1,800,000 km2). The intent of the cluster analysis is to characterize the macroscale patterns of the integrated freshwater landscape that includes lakes, wetlands, and streams and their surface connectivity attributes. We define freshwater connectivity as the permanent surface hydrologic connections that link lakes, wetlands, and streams and measure connectivity as the landscape position of systems within stream networks. Geographic data used in the analysis are in LAGOS-NE-GEO database v. 1.03 (Lake multi-scaled geospatial and temporal database), an integrated, multi-thematic geographic database (Soranno et al. 2015). The integrated freshwater clusters were created through a multi-step process as follows: 1) we quantified multiple freshwater connectivity metrics for lakes, streams, and wetlands separately, 2) we performed principal components analysis (PCA) on the connectivity metric values for each freshwater type to reduce collinearity, and 3) we performed k-means cluster analysis to group spatial units with similar freshwater connectivity characteristics. The resulting freshwater clusters are representations of the macroscale patterns of freshwater abundance and connectivity in the landscape.
Freshwater connectivity clusters for lakes, wetlands, and streams at the Hydrologic Unit 12 scale in the Midwest and Northeast U.S.A. – freshwater metric variables and K-means cluster assignment
This dataset includes freshwater connectivity cluster output and principal component scores for lakes, wetlands, and streams measured at the Hydrologic Unit 12 (HU12) scale in 17 U.S. states in the Midwest and Northeast regions (appr. 1,800,000 km2). The intent of the cluster analysis is to characterize the macroscale patterns of freshwater connectivity attributes. We define freshwater connectivity as the permanent surface hydrologic connections that link lakes, wetlands, and streams and measure connectivity as the landscape position of systems within stream networks. Geographic data used in the analysis are in LAGOS-NE-GEO database v. 1.03 (Lake multi-scaled geospatial and temporal database), an integrated, multi-thematic geographic database (Soranno et al. 2015). Freshwater connectivity clusters were created separately for lakes, wetlands, and streams through a multi-step process as follows: 1) we quantified multiple freshwater connectivity metrics, 2) we performed principal components analysis (PCA) on the connectivity metric values for each freshwater type to reduce collinearity, and 3) we performed k-means cluster analysis to group spatial units with similar freshwater connectivity characteristics. The resulting freshwater clusters are representations of the macroscale patterns of lake, wetland, and stream connectivity in the landscape.
Data bundle for "Spherical-angular dark field imaging and sensitive microstructural phase clustering with unsupervised machine learning"
<p>Prepared by Tom McAuliffe (t.mcauliffe17@imperial.ac.uk)</p> <p>This repository is a release of the raw data and analysis results for: 'Spherical-angular dark field imaging and sensitive microstructural phase clustering with unsupervised machine learning' </p> <p>The raw data is given as 'yprime.h5' - this contains patterns and metadata in the Bruker-exported format.</p> <p>Scripts for dataset decomposition into latent factors are given in 'Scripts'.</p> <p>Our spherical analysis code is included in 'SphericalAngleDF'.</p> <p>Outputs of our analysis code are contained in 'Analysis'.</p> <p>Figures for the paper are included in 'Figures'.</p> <p> </p>
Characteristics of human and viral RNA binding sites and site clusters recognized by SRSF1 and RNPS1
<p>This dataset was developed for the following article:</p> <p> Rogan PK, Mucaki EJ and Shirley BC. A proposed molecular mechanism for pathogenesis of severe RNA-viral pulmonary infections [version 1; peer review: awaiting peer review]. <em>F1000Research</em> 2020, <strong>9</strong>:943 (<a href="https://doi.org/10.12688/f1000research.25390.1">https://doi.org/10.12688/f1000research.25390.1</a>)</p> <p><strong>Section 1. Extended Data Tables</strong></p> <p>This archive contains the extended data tables for the research article "A proposed mechanism for molecular pathogenesis of severe RNA-viral pulmonary infections". These tables provide SRSF1, RNPS1 and hnRNP A1 binding site and information-dense cluster counts across various RNA viral genomes [including multiple SARS-CoV-2 and influenza strains] and the human transcriptome, the estimated SARS-CoV-2 doubling time necessary for viral genome SRSF1 binding site availability to exceed sites within the host transcriptome, and an analysis of influenza, dengue, and aplastic anemia patients misdiagnosed as irradiated by established radiation gene signatures.These tables are:</p> <p><strong>Section 1 - Table 1.</strong> RNPS1 and hnRNPA1 binding sites and Information-Dense Clusters for RNPS1 and<br> hnRNPA1 in RNA Virus Genomes<br> <strong>Section 1 - Table 2A.</strong> Detailed Analysis of Information-Dense Clusters for SRSF1 (Replicate 1) in RNA Virus<br> Genomes<br> <strong>Section 1 - Table 2B.</strong> Detailed Analysis of Information-Dense Clusters for SRSF1 (Replicate 2) in RNA Virus<br> Genomes<br> <strong>Section 1 - Table 2C.</strong> Detailed Analysis of Information-Dense Clusters for RNPS1 in RNA Virus Genomes<br> <strong>Section 1 - Table 2D.</strong> Detailed Analysis of Information-Dense Clusters for hnRNP A1 in RNA Virus<br> Genomes<br> <strong>Section 1 - Table 3.</strong> Binding Site Analysis of Multiple Coronavirus Strains (Both Strands)<br> <strong>Section 1 - Table 4A.</strong> Binding Site Analysis of Multiple Influenza A (H3N2) Strains (Negative Strand Only)<br> <strong>Section 1 - Table 4B.</strong> Binding Site Analysis of Multiple Influenza A (H3N2) Strains (Both Strands)<br> <strong>Section 1 - Table 5.</strong> SRSF1, RNPS1 and hnRNPA1 Binding Sites and Information-Dense Clusters by Gene<br> <strong>Section 1 - Table 6A.</strong> Transcriptome-Wide Information Dense Clusters Intersecting DRIP- and DRIPc-seq<br> Intervals<br> <strong>Section 1 - Table 6B. </strong>Exome-Wide Information Dense Clusters within DRIP- and DRIPc-seq Intervals<br> <strong>Section 1 - Table 6C.</strong> Transcriptome-Wide Scan of Strong Binding Sites Intersecting DRIP- and DRIPc-seq<br> Intervals<br> <strong>Section 1 - Table 6D. </strong>Exome-Wide Scan of Strong Binding Sites within DRIP- and DRIPc-seq Intervals<br> <strong>Section 1 - Table 7.</strong> Rate of False Positives for Influenza, Dengue Virus and Aplastic Anemia Using<br> Radiation Signatures<br> <strong>Section 1 - Table 8.</strong> Radiation Model Genes Contributing to False Positives for Patients with Influenza A,<br> Dengue Virus, and Aplastic Anemia<br> <strong>Section 1 - Table 9A.</strong> Doubling Time of SARS-CoV-2 Needed to Exceed Host Transcriptome SRSF1 Binding<br> Sites (Positive-Strand Sites Only)<br> <strong>Section 1 - Table 9B.</strong> Doubling Time of SARS-CoV-2 Needed to Exceed Host Transcriptome SRSF1 Binding<br> Sites (Both Strands Considered)</p> <p><strong>Section 2. All SRSF1, hnRNPA1 and RNPS1 binding site tracks for human and viral genomes</strong></p> <p>We provide bedgraph tracks which provide the location and strength of binding sites (and binding site clusters) for SRSF1, RNPS1 and hnRNPA1 across the human transcriptome (GRCh37), the human exome (including +/-300nt surrounding the exon; non-intergenic only), and for all viral genome investigated in this study (Coronavirus, Dengue, HIV-1 [two strains] and Influenza [two strains]). Note that if no clusters were found for a particular viral genome, a file for said genome will not be present in the Zenodo archive.</p> <p>Folder “Cluster-to-DRIPseq-Intersection-Tracks” contain tracks which indicate where binding site clusters have been identified, intersected with DRIP-seq and DRIPc-seq intervals which indicate where there is evidence of R-Loop formation in the human genome. The DRIP-seq dataset (GSE68845) is not strand specific. DRIPc-seq (GSE70189) is strand specific, and has been taken into account in the intersection (e.g. tracks only list positive strand clusters found in positive-strand DRIPc-seq intervals).</p> <p>Due to sheer size, the human transcriptome and exome tracks which indicate the location of individual binding sites are split into two separate files (separated by strand). While the custom tracks containing human binding site information are designed to be uploaded to the UCSC Genome Browser, files containing transcriptome-wide binding site information may be too large to be uploaded and may require further filtering (i.e. by chromosome).</p> <p>To be classified as a cluster, binding sites on the same strand must have <em>Ri</em> values which sum to >50 bits, each binding site must have a neighboring site within 25nt, and all binding sites in the cluster must have <em>R<sub>i</sub></em> greater than a minimum bit threshold. For human transcriptomes and exomes, this bit minimum was set to <em>R<sub>sequence</sub></em>. The bit minimum for viral binding sites was set to 0.1 * <em>R<sub>sequence</sub></em>. The information density-based clustering algorithm utilized in this work is described in Lu and Rogan 2018 (<a href="https://f1000research.com/articles/7-1933/v2">https://f1000research.com/articles/7-1933/v2</a>) and archived source code is available through Zenodo (<a href="https://dx.doi.org/10.5281/zenodo.1892051">https://dx.doi.org/10.5281/zenodo.1892051</a>).</p> <p><strong>Section 3. Binding site clusters - lollipop plots</strong></p> <p>Lollipop plots present the genomic coordinates and information densities of clusters across the human transcriptome, human exome, and viral genomes (Coronavirus, Dengue, HIV-1 [two strains] and Influenza [one strain]). The height of the "lollipop" corresponds to the information density of a cluster. Labels above "lollipops" present the start and end genomic coordinate (GRCh37) of the cluster followed by the number of sites in the cluster enclosed in brackets. Lollipop plots associated with human transcriptomes/exomes each contain a single gene. Influenza has 8 segments and each segment requires its own plot, other viral genomes examined are presented in a single plot.</p> <p>File naming convention for human plots:</p> <ul> <li>RBP_Gene.png</li> <li>e.g. RNPS1_ADK.png</li> </ul> <p>File naming convention for viral plots (elements in square brackets do not always appear):</p> <ul> <li>Virus[.InfluenzaSegment].RiThreshold.Strand.RBP.png</li> <li>e.g. Wuhan-Hu-1.complete-genome.4.2-bits.PosStrand.hnRNPA1.png</li> </ul> <p>The specified Ri threshold indicates all binding sites which comprise a cluster have <em>R<sub>i</sub></em> greater-than or equal to the threshold.</p> <p><strong>Section 4. Ri(b,l) matrices for all binding sites scanned</strong></p> <p>The information theory-based position weight matrices for the following RNA binding proteins (RBP) used in this study: SRSF1, hnRNPA1 and RNPS1. We investigated binding using two different RNPS1 binding models. While similar, these two models contained binding site information on opposing sides of the binding site motif which is why we found it prudent to scan with both models.</p> <p>Structure of each file:</p> <p>Line #1: Start position, End position and<em> R<sub>sequence</sub></em> [average strength of sequences used to generate the model]</p> <p>Subsequent lines describe the information on each position of the binding site:</p> <ul> <li>First four columns: <em>R<sub>i</sub></em> contribution of nucleotide at this position of the matrix [A, C, G, T]</li> <li>Row #5: Position of the matrix</li> <li>Last four columns: Number of binding sites used to generate model with a particular nucleotide at this position of the matrix [A, C, G, T]</li> </ul> <p>Example:</p> <p>-2.965775 1.282153 0.034225 -4.906891 0 1 19 8 0</p> <p>At zero position of the matrix (first nucleotide), a ‘C’ would have a positive contribution to binding site strength, a ‘G’ would be relatively neutral, and an ‘A’ or ‘T’ would negatively contribute to binding site strength.</p> <p>Generation of R<sub>i</sub>(b,l) matrices and computation of <em>R<sub>i</sub></em> values and can be accomplished by utilizing the Delila package (<a href="https://alum.mit.edu/www/toms/delila/delilaprograms.html">https://alum.mit.edu/www/toms/delila/delilaprograms.html</a>).</p> <p><strong>Section 5. Ri and intersite distance - histograms</strong></p> <p>Two sets of histograms present <em>R<sub>i</sub></em> distribution and intersite distance distribution across the human transcriptome, human exome, and viral genomes (Coronavirus, Dengue, HIV-1 [two strains] and Influenza [one strain]). </p> <p>File naming convention for human plots (elements in square brackets do not always appear):</p> <ul> <li>[IntersiteDistancesThreshold-]Human-[DRIPc]-AllChrs-RBP[-RiThreshold].png</li> <li>e.g. IntersiteDistances500-Human-AllChrs-hnRNPA1-4.6-bits.png</li> </ul> <p>File naming convention for viral plots (elements in square brackets do not always appear):</p> <ul> <li>[IntersiteDistancesThreshold-]Strand-RBP-Virus[.InfluenzaSegment][-RiThreshold].png</li> <li>e.g. IntersideDistances1000-PosStrandOnly-SRSF1-top50000sitesReplicate1-HIV-1-Strain-B.png</li> </ul> <p>Intersite distance thresholds of 500 or 1000 were assigned for all intersite distance histograms. Any distances above the corresponding threshold were excluded from the plot. Plots presenting <em>R<sub>i</sub></em> distributions contain a dashed line indicating <em>R<sub>sequence</sub></em> if it is visible within the scope of the plot.</p> <p><strong>Section 6. Perl Scripts and Descriptions</strong></p> <p>This archive contains all Perl scripts discussed in this archive's associated manuscript and a document file which describes them ("Perl-Script-Descriptions-Page.docx"). The programs and their general functions are as follows:</p> <p>“ClusterToDRIPseqAnalysisProgram.pl” – reports which information-dense clusters are located within DRIPc- and/or DRIP-seq intervals (individually and by gene)</p> <p>“ClusterToDRIPseqAnalysisProgram.GeneDensityFinder.pl” – uses the output from script “ClusterToDRIPseqAnalysisProgram.pl” to determine the number and the density of information-dense clusters within a gene (total clusters within the gene and those within DRIPc-seq intervals)</p> <p>“calculateIntersiteDistance.pl” – determines the distance between all binding sites in the same gene from a list of genomic coordinates</p> <p>“removeOutliersHigherThanN.pl” – discards intersite distances computed by script “calculateIntersiteDistance.pl” that are greater than a specified threshold</p> <p>“getStatisticsOnCol.pl” – calculates the count, geometric mean, median, arithmetic mean, and standard deviation of values from the output of script “removeOutliersHigherThanN.pl”</p> <p>“ScanDataSummaryProgram.pl” – determines the number of binding sites (above a specified <em>R<sub>i</sub></em> threshold) found within known genes (the program also reports the total expression of those genes using external A549 and pneumocyte expression datasets) from binding site coordinate data</p> <p>“TotalBindingSitePerCellCalculator.pl” – estimates the number of binding sites expressed in a single A549 or pneumocyte cell at any given time.</p>
Data for "Atomic structure of solute clusters in Al-Zn-Mg alloys"
<p>This dataset contains the data used in the publication entitled "<a href="https://www.sciencedirect.com/science/article/abs/pii/S1359645420310119"><strong>Atomic structure of solute clusters in Al-Zn-Mg alloys</strong></a>", published in Acta Materialia 17. December 2020.</p> <p>The data contained herein are:</p> <ul> <li>As-acquired transmission electron microscopy (TEM) images.</li> <li>Atom probe tomography data.</li> <li>All structural models used in density functional theory (DFT) calculations.</li> <li>Structures used for simulating scanning-TEM (STEM) images and nanobeam diffraction (NBD) patterns.</li> </ul> <p> </p> <p>The TEM images includes high angle annular dark field (HAADF) images and selected area diffraction patterns. These are given in .dm3/.dm4 files, and can be opened in e.g. the "<a href="https://www.gatan.com/products/tem-analysis/gatan-microscopy-suite-software">Gatan Microscopy Suite" </a>software. The images are also given as .tif images. The files are names after the "Figx_alloy_condition_xxx". "Figx" refers to the figure in the main article, "alloy" describes the alloy used and "condition" describes from what ageing condition. The uncorrected image series used for Fig. 6c (in the article) is included and requires the <a href="http://lewysjones.com/software/smart-align/">SmartAlign </a>plugin in the Gatan Microscopy Suite to analyse the dataset. SmartAlign allows for correcting rigid and non-rigid distortions in the STEM images in order to reduce effect of specimen drift and scan noise during acquisition. </p> <p>The ATP data is given as a .xlsx file. The data here is the processed data after applying the maximum separation algorithm. The data here is used to produce Figs. 2b and 2c in the paper. <br> <br> The structures used in the DFT calculations are given here as .cif files. These are separated into "Single_clusters" and "Stacked_clusters" and named according to Tabs. 1 and 2 in the Supplementary material of the paper.</p> <p>The two structures used for simulating STEM-HAADF and NBD patterns are given in the folder "TEM_simulations". "Mg32Zn124D_94x94" was used for NBD and "Mg32Zn124D_X_Zn4" was used for HAADF-STEM. The stack used for Supplementary Fig. 7c is labeled "Mg32Zn124D_94x94_slab_1Allayerop.cif".</p> <p> </p> <p> </p> <p> </p>
North Atlantic jet stream clusters: daily and seasonal occurence
<p>This dataset contains the time series used in Madonna et al 2020 (Reconstructing winter climate anomalies in the Euro-Atlantic sector using circulation patterns, DOI: 10.5194/wcd-2021-6)</p> <p><br> Filenames:</p> <p>1) seasonal_timeseries.txt</p> <p>Time series of the occurrence (in % = days/season*100) of time during winter of each jet cluster, blocking and the NAO.<br> Winters are defined as December, January and February (DJF). The season name is given by the last month (i.e. 1980 is December 1979, January 1980 and February 1980). 29 February is removed from the data so that each winter season has 90 days.</p> <p>Jet clusters are calculated following Madonna et al 2017. The five clusters are named as in Madonna et al 2017: Northern (N), Central (C), Mixed (M), Southern (S) and Tilted (T).<br> Blocking are calculated following Scherrer et al. 2006 and averaged over Greenland (GB), offshore of the Iberian Peninsula also called Iberian wave breaking (IWB) and over Scandinavia (SBL). The exact definition of the regions can be found in Madonna et al 2020.</p> <p>The NAO index was downloaded from ftp://ftp.cpc.ncep.noaa.gov/cwlinks/norm.daily.nao.index.b500101.current.ascii. Positive (NAO+) and negative (NAO-) days are defined as those that exceed 0.5 DJF standard deviation, corresponding to values greater than 0.613 and lower than -0.177, respectively.</p> <p>Example: during winter 1980, 7.78% of the days were in the North jet cluster. This is equivalent to 7 days -> 7.78 * 90 (days per season) /100</p> <p><br> 2) daily_inverse_distance_from_centroid.txt contains information about the similarity of the 2D zonal wind field to the cluster centroids which is used to determine the jet state.</p> <p>The file has 12 columns, labelled as follow:<br> date, lat, speed, N4, C4, M4, S4, N5, C5, M5, S5, T5</p> <p>The first column (date) shows the day in YYYYMMDD format, the second (lat) is the latitude (in °N) of the maximum zonal wind in the 60°W-0°W sector (i.e. the jet latitude index, see Woollings et al. 2010 or Madonna et al. 2017 for more details), and the third (speed) is the zonal averaged (60°W-0°) zonal wind speed (in m/s) at the latitude given by column 2.</p> <p>Columns 4-7 give the inverse distance from each cluster centroids using four (4) clusters: Northern (N4), Central (C4), Mixed (M4), Southern (S4) and is normalized from 0 to 1. Values close to 1 means that the clusters are similar to its centroid. The distances sum up to 1.</p> <p>Columns 8-12 show similar to 4-7 the inverse distance from the centroids using five (5) clusters: Northern (N5), Central (C5), Mixed (M5), Southern (S5) and Tilted (T5). Distances are also normalized and sum up to 1.</p> <p>In the study of Madonna et al 2020, a day has a defined cluster X (X=N, C, M, S, T), if the inverse distance from the cluster centroid X exceeds 0.5 and it clearly dominates over the other clusters.</p> <p><br> Example: 1 January 1979, the zonal mean zonal wind is maximum at 47°N and has a value of 15.61 m/s.<br> Considering 4 clusters, the jet resembles most the Mixed cluster (M4=0.36), followed by the Southern (S4=0.23), Northern (N4=0.22) and Central (C4=0.18). The sum of the distances (0.36 + 0.23 + 0.22 + 0.18 = 0.99 due to decimal approximation) is equal to 1. Using 4 clusters, this day would be assigned to cluster M4. The day is, however, not clearly identified as a Mixed jet, as the inverse distance (M4=0.36) is smaller than 0.5. The threshold of 0.5 is set to identify days where a centroid clearly leads over the others.<br> If we consider 5 clusters, the jet on 1 Jan 1979 resembles the tilted jet (T5 = 0.73) and has very little in common with the other centroids (values of 0.05-0.08). Thus, considering 5 clusters, this day is classified as a tilted jet. It is also clearly defined, as 0.73 > 0.5.</p> <p><br> References:</p> <p>Madonna, E., Li, C., Grams, C.M. and Woollings, T. (2017), The link between eddy‐driven jet variability and weather regimes in the North Atlantic‐European sector. Q.J.R. Meteorol. Soc, 143: 2960-2972. https://doi.org/10.1002/qj.3155</p> <p>Scherrer, S. C., Croci‐Maspoli, M., Schwierz, C., and Appenzeller, C. (2006). Two‐dimensional indices of atmospheric blocking and their statistical relationship with winter climate patterns in the Euro‐Atlantic region. International Journal of Climatology, 26(2), 233-249</p> <p>Woollings T, Hannachi A and Hoskins B. (2010). Variability of the North Atlantic eddy‐driven jet stream. Q. J. R. Meteorol. Soc. 136: 856– 868.</p> <p> </p> <p> </p> <p> </p>
Physioclimatic clusters of Austria
<p><strong>physioclimatic_features_grid_AT_average_1992-2021.nc</strong></p> <p>A netcdf dataset on a 1 km grid covering Austria and it's associated catchment areas. The spatial dimensions are 329 by 584 gridpoints in the projection ETRS89 / Austria Lambert (<a href="https://epsg.io/3416">EPSG:3416</a>). The data consist of 176 different climatological and geomorphometric indices, which are calculated for each gridpoint and then averaged across the climatological normal 01-01-1992 to 31-12-2021. The basis variables from which indices are calculated are elevation, temperature, precipitation, reference evapotranspiration, sunshine duration, snow height and snow water equivalent.</p> <p> </p> <p><strong>physioclimatic_clusters_raster_AT.tif</strong></p> <p>A GeoTiff file of derived physioclimatic subregions comprising 7 characteristic clusters and one noise class, which are the main regionalisation results of the paper below.</p> <p> </p> <p><strong>physioclimatic_clusters_vector_AT.gpkg</strong></p> <p>A GeoPackage consisting of vectorized features of type <em>multipolygon </em>comprising spatially filtered physioclimatic regions based on the gridded output clusters. Note that the small valley clusters have been filtered out and the remaining larger clusters have been spatially joined in order to derive simple multipolygons. This derived, spatially filtered set of clusters can be used for less granular applications compared to the above fine-grained cluster output.</p> <p> </p> <p><strong>base_variables.nc</strong></p> <p>Base Variables mean Temperature, Precipitation and slope, used to evaluate the clustering output.</p> <p> </p> <p><strong>principal_components_dim20.nc</strong></p> <p>First 20 principal components of the full feature space, used as direct input for UMAP/HDBSCAN and k-means and to evaluate the clustering output.</p> <p> </p> <p>Please refer to the <a href="https://github.com/Geosphere-Austria/subregion-derivation">GitHub</a> repository for further details. The peer-reviewed, open access paper can be found here: <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.envsoft.2025.106324" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.envsoft.2025.106324</a>.</p> <p> </p> <p><strong>Version History</strong></p> <p>v1.0.1: Fixed a small hole between two polygons in the southern part of the domain in the physioclimatic_clusters_vector_AT.gpkg file.</p> <p>v1.0.0: Initial upload.</p>
Dataset for Investigating Anomalies in Compute Clusters
<p><strong>Abstract</strong></p><p>The dataset was collected for 332 compute nodes throughout May 19 - 23, 2023. May 19 - 22 characterizes normal compute cluster behavior, while May 23 includes an anomalous event. The dataset includes eight CPU, 11 disk, 47 memory, and 22 Slurm metrics. It represents five distinct hardware configurations and contains over one million records, totaling more than 180GB of raw data.</p><p><strong>Background</strong></p><p>Motivated by the goal to develop a digital twin of a compute cluster, the dataset was collected using a Prometheus server (1) scraping the Thomas Jefferson National Accelerator Facility (JLab) batch cluster used to run an assortment of physics analysis and simulation jobs, where analysis workloads leverage data generated from the laboratory's electron accelerator, and simulation workloads generate large amounts of flat data that is then carved to verify amplitudes. Metrics were scraped from the cluster throughout May 19 - 23, 2023. Data from May 19 to May 22 primarily reflected normal system behavior, while May 23, 2023, recorded a notable anomaly. This anomaly was severe enough to necessitate intervention by JLab IT Operations staff.</p><p>The metrics were collected from CPU, disk, memory, and Slurm. Metrics related to CPU, disk, and memory provide insights into the status of individual compute nodes. Furthermore, Slurm metrics collected from the network have the capability to detect anomalies that may propagate to compute nodes executing the same job.</p><p><strong>Usage Notes</strong></p><p>While the data from May 19 - 22 characterizes normal compute cluster behavior, and May 23 includes anomalous observations, the dataset cannot be considered labeled data. The set of nodes and the exact start and end time affected nodes demonstrate abnormal effects are unclear. Thus, the dataset could be used to develop unsupervised machine-learning algorithms to detect anomalous events in a batch cluster.</p><p><a href="https://doi.org/10.48550/arXiv.2311.16129">https://doi.org/10.48550/arXiv.2311.16129</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.