Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,326
datasets available to search
ShareScore release 0.9.0
Dataset results
2,326 results for “clusters”
Metaclusters by DPCfam clustering of UniRef50 v 2017_07
<p>Metaclusters obtained from the DPCfam clustering of UniRef50, v. 2017_07.<br> Metaclusters represent putative protein families automatically derived using the DPCfam method, as described in <em>Unsupervised protein family classification by Density Peak clustering, Russo ET, 2020, PhD Thesis <a href="http://hdl.handle.net/20.500.11767/116345">http://hdl.handle.net/20.500.11767/116345</a> . Supervisors: Alessandro Laio, Marco Punta.</em></p> <p>Visit also <a href="https://dpcfam.areasciencepark.it/">https://dpcfam.areasciencepark.it/</a> to easily navigate the data.</p> <p><strong>VERSION 1.1 changes:</strong></p> <ul> <li>Added DPCfamB database, including all small metaclusters with 25<=N<50 seed sequences. DPCdamB files are named with the prefix B_</li> <li>Added Alphafold representative based on AlphaFoldDB for each MC</li> </ul> <p><strong>FILES DESCRIPTION:</strong></p> <p><strong>1) Standard DPCfam database</strong></p> <ul> <li><strong>metaclusters_xml.tar.gz </strong>Metaclusters' seeds, unaligned in an xml table. Only MCs with seeds with 1) more than 50 elements and 2) average length larger than 50 a.a.s are reported. Metaclusters entries include also some statistical information about each MC (such as size, average length, low complexity fraction etc, ) and Pfam comparison (Dominant Architecture). A README file is included describing the data. A parser is included to transform XML data to space-separated tables. XML schema is included.</li> <li><strong>metaclusters_msas.tar.gz</strong> Metsclusters' multiple sequence alignments, in fasta format. Only MCs with seeds with 1) more than 50 elements and 2) average length larger than 50 a.a.s are reported .</li> <li><strong>metaclusters_hmms.tar.gz</strong> Metsclusters' profile-hmms. A ".hmm" file for each metacluser. Only MCs with seeds with 1) more than 50 elements and 2) average length larger than 50 a.a.s are reported .</li> <li><strong>all_metaclusters_hmm.tar.gz</strong> Collctive metaclusters' profile-hmm. A single .hmm file collecting all MC's profile-hmm. . Only MCs with seeds with 1) more than 50 elements and 2) average length larger than 50 a.a.s are reported </li> <li><strong>uniref50_annotated.xml.gz</strong> UniRef50 v.2017_07 database annotated with Pfam families and DPCfam metaclusters. A README file is included describing the data. A parser is included to transform XML data to space-separated tables. XML schema is included. XML schema is derived from uniprot's UniRef50 xml schema.</li> </ul> <p><strong>2) DPCfamB database</strong></p> <ul> <li><strong>B_metaclusters_xml.tar.gz </strong>Metaclusters' seeds, unaligned in an xml table. All metaclusters are listed. Metaclusters entries include also some statistical information about each MC (such as size, average length, low complexity fraction etc, ) and Pfam comparison (Dominant Architecture). A README file is included describing the data. A parser is included to transform XML data to space-separated tables. XML schema is included. </li> <li><strong>B_metaclusters_msas.tar.gz</strong> Metsclusters' multiple sequence alignments, in fasta format. Only MCs with seeds with 1) 25<=N<50 elements and 2) average length larger than 50 a.a.s are reported .</li> <li><strong>B_metaclusters_hmms.tar.gz</strong> Metsclusters' profile-hmms. A ".hmm" file for each metacluser. Only MCs with seeds with 1) 25<=N<50 elements and 2) average length larger than 50 a.a.s are reported .</li> <li><strong>B_ all_metaclusters_hmm.tar.gz</strong> Collctive metaclusters' profile-hmm. A single .hmm file collecting all MC's profile-hmm. . Only MCs with seeds with 1) 25<=N<50 elements and 2) average length larger than 50 a.a.s are reported </li> </ul> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
Beyond bold versus shy: Zebrafish exploratory behavior falls into several behavioral clusters and is influenced by strain and sex
<p>Individual differences in exploratory behavior have been found across a range of taxa and are thought to contribute to evolutionary fitness. Animals that explore more of a novel environment and visit areas of high predation risk are considered bold, whereas animals with the opposite behavioral pattern are shy. Here, we determined whether this bimodal characterization of bold versus shy adequately captures the breadth of behavioral variation in zebrafish or if there are more than these two subtypes. To identify behavioral categories, we applied unsupervised machine to three-dimensional swim traces from over 400 adult zebrafish across four strains (AB, TL, TU, and WIK) and both sexes. We found that behavior stratified into four distinct clusters: previously described bold and shy behavior and two new behavioral types we call wall-huggers and active explorers. Clusters were stable across time and influenced by strain and sex where we found that TLs were shy, female TU fish were bold, male TU fish were active explorers, and male ABs were wall-huggers. Our work suggests that zebrafish exploratory behavior has greater complexity than previously recognized and lays the groundwork for the use of zebrafish in understanding the biological basis of individual differences in behavior.</p>
Semi-empirical error ellipsoid clustering for identifying the second-order structural features from a laboratory AE source location cloud—method, validation, and application to a hydraulic fracturing test [DATA]
<p>Data and metadata for the publication "Semi-empirical error ellipsoid clustering for identifying the second-order structural features from a laboratory AE source location cloud—method, validation, and application to a hydraulic fracturing test", published in Earth and Space Science.</p>
Variable allelic expression of imprinted genes at the Peg13, Trappc9, Ago2 cluster in single neural cells
<p>Fig 1 5' RACE - Sequence tracks and analysis of alternative Trappc9 transcriptional start sites</p> <p>Fig 2 and suppl fig 3- Brain and Kidney tissue or Neural stem cells isolated from the Hippocampus region of newborn mice generated from a C57BL/6 (female) and a Cast/EiJ (male) cross (and its reciprocal cross) was used to determine allelic bias expression. A SNP located within an exon of Peg13, Trappc9, Ago2, Chrac1 and Kcnk9 was identified and amplified via pyrosequencing PCR with the percentage of SNP identification used to determine allele expression percentages. Additionally, Pyrorun sequences of reverse transcribed RNA from Trappc9 expression in different tissues. A hybrid cross between C57BL/6 and JF1 mouse was used to generate hybrid pups that were used to determine allele specificity of Trappc9 expression in Kidney and brain tissues.</p> <p>Figs 3, 4 & 5- Single neural stem cells were isolated from the Hippocampus of newborn mice generated from a hybrid cross. Some of these cells were differentiated In vitro and either the NSC or differentiated neurons were lysed and underwent a reverse transcription. The newly formed cDNA was used as a template to amplify expressed Peg13, Trappc9 or Ago2 transcripts which was then sent for Sanger sequencing. A SNP located within the exon was used to determine whether the transcript from that cell was generated from the maternal or paternal allele.</p> <p>Fig 6- Brain-specific regulatory elements were cloned into a pGL [Luc] vector containing a Trappc9 promoter. These newly generated plasmids were then transfected into either primary neuron or fibroblast cultures alongside a Renilla vector for normalization using Lipofectamine as a transfection reagent. After 48 hours the cells were lysed and analyzed using a Glomax illuminator to determine their impact on Luciferase expression compared to that of the pGL vector containing just the Trappc9 promoter. Additionally, Plasmid vectors that were used for transfection of primary neurons to determine the impact of brain-specific regulatory elements on transcription. Plasmids can be visualized using the free software pdraw.32 downloaded from http://acaclone.com/download/install.htm</p> <p>Suppl fig 2- Single Neural stem cells and in vitro differentiated neurons isolated from Mouse Hippocampus tissue underwent reverse transcription and a qPCR reaction intended to amplify cDNA of genes associated with specific neural cell types as a method of detecting cell fate and whether this had an impact on allele-specific expression and differential methylation.</p> <p>Suppl fig 4 & 5- DNA isolated from Neural stem cells was bisulfite converted for downstream identification of methyl group presence. CpG islands located at or near the promoters of the Peg13, Trappc9, Ago2, Chrac1 and Kcnk9 genes were amplified and cloned into a TOPO vector. The cloned segments were Sanger sequenced and compared to the original non-bisulfite converted sequence using the Quantification for methylation analysis (QUMA) tool to determine which CG dinucleotides were methylated and which weren't. Pyroruns determining methylation frequency at the CpG islands of these genes can be found in a separate upload on Zenodo.</p>
Replication Package for ICSE 2023 submission "How Deep Learning Packages Form Supply Chain in PyPI: Types, Clusters, and Detachment"
<p>This is the replication package for our ICSE 2023 submission <em><strong>How Deep Learning Packages Form Supply Chain in PyPI: Types, Clusters, and Detachment</strong></em>. </p>
Companion data for Communication-Aware Load Balancing of the LU Factorization over Heterogeneous Clusters
<p>This is the companion data repository for the paper entitled <strong>Communication-Aware Load Balancing of the LU Factorization over Heterogeneous Clusters</strong> by Lucas Leandro Nesi, Lucas Mello Schnorr, and Arnaud Legrand. The manuscript has been accepted in the <a href="https://icpads2020.comp.polyu.edu.hk/">ICPADS 2020</a>.</p>
Companion data of Detection, Evaluation and Mitigation of Resource Affinity and Communication Contention Problems in a Task-Based Runtime over Heterogeneous Clusters
<p>This is the companion data repository for the paper entitled <strong>Detection, Evaluation, and Mitigation of Resource Affinity and Communication Contention Problems in a Task-Based Runtime over Heterogeneous Clusters</strong> by Lucas Leandro Nesi and Lucas Mello Schnorr. The manuscript has been accepted for publication in the <a href="http://wscad.sbc.org.br/current/index.html">WSCAD 2020</a>.</p>
Data Set "Accurate quantum-chemical fragmentation calculations for ion–water clusters with the density-based many-body expansion"
<p>This data set accompanies the publication "Accurate quantum-chemical fragmentation calculations for ion–water clusters with the density-based many-body expansion"</p> <p>It contains:</p> <p>- xyz files of all considered molecular structures.</p> <p>- PyADF input scripts for running the eb-MBE and db-MBE calculations.</p> <p>- raw results data from the eb-MBE and db-MBE calculations</p> <p>- Jupyter notebooks for generating the plots and tables</p>
A model for the dissemination of circulating tumour cell clusters involving platelet recruitment and a plastic switch between cooperative and individual behaviours
<p>This folder includes live/dead cell counts as well as transwell assay data for the corresponding manuscript. </p>
Clustering of psychological profiles of primary school pupils
<p>Traditionally, arbitrary psychological research on different types of student activities uses certain sets of tests that reflect different states of students. These states are defined on the basis of certain criteria, the values of which are usually reflected by numerical scales with different intervals. Table 1 provides a list of these profiles and their factors, evaluation criteria with corresponding scales. Table 2 presents generalized results of testing of primary school pupils of the 2nd grade. <br> Due to the fact that the whole study is implemented on the basis of certain psychological profiles, which we also interpret as hyperproperties, in our case they are the receptors. <br> So the psychological profile of school motivation as a hyper-property defines five receptors, namely: <br> school motivation_high level <br> school motivation_good school motivation <br> school motivation_positive attitude to school<br> school motivation_low school motivation <br> school motivation_negative attitude to school<br> The protocol of cluster formation with receptor names is given in the file result of clustering.rtf</p> <p>Традиційно довільне психологічне дослідження різних видів діяльності учнів використовує певні набори тестів, які відображають різні стани учнів. Ці стани визначаються на основі певних критеріїв, значення яких звичайно відображаються числовими шкалами з різними інтервалами. вказані критерії оцінювання можуть бути проінтерпретовані як відповідні властивості учнів, що мають різні рівні прояву в їх діяльності. У табл. 1 наведено перелік цих профілів та їх фактори, означено критерії їх оцінювання з відповідними шкалами. У табл.2 наведені узагальнені результати тестування молодших учнів 2-го класу. <br> У зв’язку з тим, що усе дослідження реалізується на основі певних психологічних профілів, які ми ще й інтерпретуємо як гіпервластивості, у нашому випадку рецепторами є саме вони. <br> Так психологічний профіль school motivation, як гіпервластивість визначає п’ять рецепторів, а саме: <br> school motivation_high level <br> school motivation_good school motivation <br> school motivation_positive attitude to school<br> school motivation_low school motivation <br> school motivation_negative attitude to school<br> Протокол утворення кластерів з іменами рецепторів наведено в файлі result of clustering.rtf</p> <p> </p> <p> </p>
Data from: Clustering Deviation Index (CDI): A robust and accurate internal measure for evaluating scRNA-seq data clustering
<div> <div> <p>The clustering of cells has been widely used to explore the heterogeneity of cell populations in single-cell RNA-sequencing (scRNA-seq). We proposed a parametric model for monoclonal and polyclonal scRNA-seq data to evaluate clustering results. Based on the parametric model, we proposed a metric (CDI) to quantify the goodness-of-fit of cell clustering to the data. Here we presented CT26.WT and T-CELL as two datasets to examine the performance of our model and metric. CT26.WT contains wild-type CT26 cells from the murine colorectal carcinoma cell line, and cells in CT26.WT are highly homogeneous. T-CELL contains T-cells from tumor tissue of mice three weeks after 4T1 tumor injection. From these datasets and public datasets, we validated our model and benchmarked our metric.</p> </div> </div>
Recent HIV infections among newly diagnosed individuals living with HIV in rural Lesotho: Secondary data from the VIBRA cluster-randomized trial
<p>These are pseudo-anonymised data from a secondary analysis of the VIBRA randomised trial.</p> <p>HIV recency assays are used to distinguish recently acquired infection from long-term infection among individuals newly diagnosed with HIV. Since 2015, the World Health Organisation recommends the use of an algorithm to assess recency of infections which is based on an HIV recency assay and viral load (VL) quantification. We determined the proportion of recent HIV infections among participants of the VIBRA (Village-Based Refill of Antiretroviral therapy) cluster-randomized trial in Lesotho and assessed risk factors for these recent infections.</p> <p>The VIBRA trial recruited individuals living with HIV and not taking antiretroviral therapy during a door-to-door HIV testing campaign in two rural districts (Butha-Buthe and Mokhotlong). Samples were collected from participants newly diagnosed and tested for HIV recency using the Asanté HIV-1 Rapid Recency Assay and VL using the Roche Cobas System. Clinical and socio-demographic data were extracted from the trial database. Univariate analysis was conducted to determine factors associated with recent compared to long-term infection.</p> <p>There is one dataset containing all data presented in this dtuy. The data codebook explains the data available in the dataset.</p>
Supporting Information for the Journal Article "Systematic microsolvation approach with a cluster-continuum scheme and conformational sampling"
<p>This dataset contains the supporting information published together with the article "Systematic microsolvation approach with a cluster-continuum scheme and conformational sampling" (<a href="https://doi.org/10.1002/jcc.26161"><em>J. Comput. Chem.</em>, <strong>2020</strong>, <em>41</em>, 1144</a>).</p>
Dataset: Quantifying cell densities and biovolumes of phytoplankton communities and functional groups using scanning flow cytometry, machine learning and unsupervised clustering
<p>This dataset contains all relevant data for the manuscript (in submission) "<em>Quantifying cell densities and biovolumes of phytoplankton communities and functional groups using scanning flow cytometry, machine learning and unsupervised clustering</em>".</p> <p>Code written to analyse this dataset (which may be adapted for other flow cytometry datasets) is found at https://zenodo.org/record/999747</p> <p>--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------</p> <p>Naming convention for raw flow cytometry data files (located in /Script 3. Generating raw data subset/input/):</p> <p>[Allparameters] _ [Year] - [Month] - [Date] [Hour] [u] [Minute] _ [Depth]</p> <p>e.g: Allparameters_2014-07-31 08u08_1.0m</p> <p>The date, time and depth indicate the location and time at which the measurement was taken.</p>
The impact of weekly fever-screening and treatment and monthly RDT testing and treatment on the infectious reservoir of malaria in Burkina Faso: results from a cluster-randomised trial
<p>The majority of malaria infections in endemic countries are asymptomatic and a source of onward transmission to mosquitoes. Malaria transmission and disease burden could be reduced by improving early detection and treatment of these infections with active screening approaches. In an 18-month cluster-randomized study in Sapone, Burkina Faso, households were enrolled and randomised to 1 of 3 arms: arm 1 - control; arm 2 - active weekly screening for febrile individuals and treatment if rapid diagnostic test (RDT) positive; or arm 3 – active weekly fever-screening (as arm 2) plus monthly RDT-testing regardless of symptoms. The primary outcome was parasite prevalence by qPCR in the end-of-study cross-sectional survey. Secondary outcomes included parasite and gametocyte prevalence and density in all three end-of-season cross-sectional surveys, incidence of infection, and the transmissibility of infections to mosquitoes. A total of<strong> </strong>906 individuals were enrolled during 2 phases. In Phase 1, 412 individuals were enrolled between August 9 and 17, 2018, and in Phase 2, 494 individuals were enrolled between January 10 and 31, 2019. In the end-of-study cross-sectional survey, malaria parasite prevalence by qPCR was statistically significantly lower in arm 3 (29·26% 79/270), but not in arm 2 (45·66% 121/265), when compared to arm 1 (48·72% 133/273) (RR = 0·65, 95%CI = 0·52 to 0·81, P=0·0001). Total parasite and gametocyte prevalence and density were also significantly lower in arm 3 in all surveys. The largest differences were seen at the end of the dry season, with gametocyte prevalence 78·38 % and transmission potential 98·20% lower in arm 3 vs arm 1. Active monthly RDT testing and treatment can reduce parasite carriage and the infectious reservoir of malaria to <2% when used during the dry season. This insight may inform approaches for malaria control and elimination.</p>
Supplementary data to "High-Resolution Raman Imaging of >300 Patient-Derived Cells from Nine Different Leukemia Subtypes: A Global Clustering Approach"
<p>Compressed ".feather" files including the entire dataset of 319 Raman maps of the same number of cells from 19 patients affected by nine distinct leukemia subtypes.<br>Raw data have been pre-processed as follows using custom software (LabVIEW, National Instruments Corp., TX): a) cosmic rays removal by singular value decomposition (SVD); b) camera offset subtraction; c) CCD response correction (intensity and etaloning) using a tungsten halogen light with known emission (Avalight-HAL, Avantes BV, NL)); d) wavenumber calibration using the zero-wavenumber laser line, toluene and argon-mercury emission (CAL-2000, Ocean Optics, Germany); e) denoising by SVD.<br>More details in the open access published article and supplementary material (10.1021/acs.analchem.4c00787).</p>
Diseasome - Finding disease association based on Phenotypic and Genotypic clustering
<p>The final processed phenotypic clustered data (output.zip) from Orphanet is also added along side the genotypic clustering data.</p>
DNS dataset for modelling homogeneous ignition processes of clustering solid particle clouds in isotropic turbulence
<h2>Abstract</h2> <p>This dataset is being published to enable the development of models for igniting and combusting solid particles in isotropic turbulence using flamelet tabulated chemistry. This dataset is generated using the forced homogeneous isotropic turbulence in order to investigate the effect of active turbulent forces on the ignition of particles, and it is used as the supplementary material for the manuscript "<em>Modeling homogeneous ignition processes of clustering </em><em>solid particle clouds in isotropic turbulence</em>", which was accepted for publication in the special issue of Fuel Journal for the proceeding of the 4th international Oxyflame workshop. The particles are chosen to have near unity Stokes numbers, which promote particle clustering to study the ignition phenomenon for particle clusters, which is common during ignition and combustion of particle clouds in large industrial solid fuel-powered burners. Since the ignition process of clustering solid particle clouds is transient, different time instances during the ignition process are presented to facilitate the modelling effort for the transient ignition behaviour of the particles. The dataset consists of gas-phase data, particle data and the most important routines required to process the dataset. </p> <p>This database is a valuable resource for users developing solid fuel ignition and combustion in turbulent conditions. </p> <p>It should be noted that this dataset is a reduced version of the full dataset in order to size limitations in the sharing platforms. More information and full dataset can be provided upon request. For more information, contact: p.farmand@itv.rwth-aachen.de</p> <h2>Technical details</h2> <p>The data provided in this dataset contains gas phase data, particle data, and some post-processing scripts for visualization of the data. The data is generated in forced homogeneous isotropic turbulence (HIT) with an initial preferential concentration of the particles in a hot atmosphere to study the impact of particle clustering on ignition. Simulations were performed within a region with the physical size of 12.8mm * 12.8mm * 12.8mm with periodic boundary conditions in all directions. The domain size is discretized with a three-dimensional cartesian mesh with a resolution of Δx = 50 μm. A forced isotropic turbulent field with Re_λ=30 and the Kolmogorov length scale η =100 microns has been chosen. The dispersed phase consists of 10,000 particles of Colombian coal with D_p= 20 microns and T0=300K and with an apparent density of 700kg/m3. Non-reactive particles are first randomly distributed in the box filled with air with 20% oxygen and an initial gas temperature of T = 1500K, which is relevant to practical PCC applications. These conditions lead to an initial Stokes number of around 5, for which a clustering behaviour in particle cloud motion is expected. The employed forced isotropic turbulence ensures maintaining the same turbulence statistics during non-reactive and reactive simulations, as summarised in the following table:</p> <table> <tbody> <tr> <td> <p><em>time [ms]</em></p> </td> <td> <p><em>Re_</em><em>λ</em></p> </td> <td> <p><em>Re_</em><em>Turb</em></p> </td> <td> <p><em>η[m]</em></p> </td> <td> <p><em>l_</em><em>t</em><em>[m]</em></p> </td> <td> <p><em>t_</em><em>η</em><em>[ms]</em></p> </td> <td> <p><em>t_</em><em>l</em><em>[ms]</em></p> </td> <td> <p><em>St</em></p> </td> </tr> <tr> <td>0</td> <td>30.7</td> <td>141.8</td> <td>1.04e-4</td> <td>4.27e-3</td> <td>4.47e-2</td> <td>5.33e-1</td> <td>6.199</td> </tr> <tr> <td>0.46</td> <td>30.9</td> <td>143.9</td> <td>9.94e-5</td> <td>4.13e-3</td> <td>4.16e-2</td> <td>4.99e-1</td> <td>6.207</td> </tr> <tr> <td>0.5</td> <td>30.5</td> <td>140.1</td> <td>9.96e-5</td> <td>4.06e-3</td> <td>4.27e-2</td> <td>5.05e-1</td> <td>6.333</td> </tr> <tr> <td>0.55</td> <td>29.4</td> <td>129.8</td> <td>1.01e-4</td> <td>3.88e-3</td> <td>4.34e-2</td> <td>4.94e-1</td> <td>6.032</td> </tr> <tr> <td>0.6</td> <td>28.5</td> <td>122.1</td> <td>1.02e-4</td> <td>3.76e-3</td> <td>4.46e-2</td> <td>4.92e-1</td> <td>5.775</td> </tr> <tr> <td>0.65</td> <td>27.7</td> <td>115.2</td> <td>1.04e-4</td> <td>3.66e-3</td> <td>4.48e-2</td> <td>4.81e-1</td> <td>5.365</td> </tr> </tbody> </table> <p><em>η and t_η are the respective Kolmogorov length and time scales, and l_t and t_l correspond to the integral length and time scales.</em></p> <p> </p> <h3>Gas phase and particle data:</h3> <p>The gas phase data has HDF5 formation, which contains selected scalars relevant to model developments. Since the ignition process is a transient process, two different time instances, t=0.5ms with the maximum number of ignited regions during the ignition process and t=0.65ms at the end of the ignition process, are chosen. Before starting the reactive simulations, a non-reactive simulation for 20.5ms was performed to form the particle clusters, and then the reactive simulation was started. Therefore, the data.out_2.100E-02.h5 corresponds to t = 0.5ms and data.out_2.115E-02.h5 corresponds to t = 0.65ms. Each HDF5 dataset has the following structure:</p> <p>Group Flow:</p> <ul> <li>U, V, W</li> </ul> <p>Group scalar: </p> <ul> <li>CO, CO2, H2O, C2H2, O2, N2, OH, progress variable (PV), temperature(T), enthalpy(h), heat capacity (Cp), density(RHO), pressure(P), mixture fraction(Z), dissipation rate, heat conductivity</li> </ul> <p>Group C_dot:</p> <ul> <li>Molar production rate of CO, CO2, O2, OH</li> </ul> <p>Group ST:</p> <ul> <li>Mass fraction rate of OH</li> </ul> <p>Group var (data required for subfilter analysis and PDF modelling):</p> <ul> <li>RHO * h</li> <li>RHO * h * h</li> <li>RHO *<em> </em>PV</li> <li>RHO *<em> </em>PV <em>* </em>PV</li> <li>RHO *<em> </em>RHO</li> <li>RHO *<em> </em>RHO <em>* </em>RHO</li> <li>RHO * T</li> <li>RHO *<em> </em>T <em>* </em>T</li> <li>RHO * Z</li> <li>RHO *<em> </em>Z <em>* </em>Z</li> </ul> <p>These gas phase data can be used to study the Eulerian field, which is impacted by the particles through the source terms, which are obtained from the Lagrangian framework. For the particle data, two CSV files containing information about each particle's position and temperature are provided. </p> <p>In each CSV file, which corresponds to the same time instance as the gas phase data, this structure can be found:</p> <ul> <li>Points_0, Points_1, Points_2: particle position x,y,z</li> <li>coal_defaultT: particle temperature</li> </ul> <h3>Scripts:</h3> <p>Different scripts are provided for a reader to enable the first-time use of the data, such as visualising the gas phase quantities, calculating the conditional mean and statistical analysis, and clustering analysis for the particles. Here is a brief information regarding the scripts:</p> <ul> <li><strong>hdf5_2D_Con_Mean_Plots.m</strong>: This script can be used for loading the HDF5 file and visualising the field, joint PDF, and joint correlation of different quantities with respect to different parameters. This script requires another module to calculate the conditional mean of the data.</li> <li><strong>Compute_ConditionalMean_histograms_2D.m: </strong>This module calculates the conditional average of a 2D array.</li> <li><strong>clustering_voronoi_3D.m</strong>: This script can analyse the particle data and calculate the clustering limit based on the Voronoi algorithm. It can also filter the clustered particles and separate different clusters using the nearest neighbour and DB-SCAN methods.</li> </ul> <p> </p> <p> </p>
Benchmark dataset for CATH hierarchical clustering tools (GeMMA/FunFHMMEr, MARC, FRAN and eMMA)
<p>Benchmark dataset for CATH SuperFamily 3.40.50.620 (HUPS).</p> <p>Contains Functional Families alignments and Hidden Markov Models generated by GeMMA/FunFHMMER, MARC, FRAN and CATH-eMMA and Python code used to assess their quality (EC purity, DOPS, Neff) and intermediate steps by the MARC and FRAN pipelines (pooling, randomisation, renaming).</p> <p>3.4.50.620_full_superfamily_sequences.fasta contains all HUPs superfamily sequences, the FunFams are a subset of these.</p> <p>all_starting_clusters_sequences.fasta contain the sequences included in the starting clusters used in the analyses.</p> <p>3.40.50.620_embedded.pt includes embeddings for the HUPs superfamily generated using the ESM2 Protein Language Model.</p> <p> </p>
Data Release Scrutinising evidence for the triggering of Active Galactic Nuclei in the outskirts of massive galaxy clusters at z~1
<p>Dataset of the paper "Scrutinising evidence for the triggering of Active Galactic Nuclei in the outskirts of massive galaxy clusters at z~1".</p> <p> </p> <p>All the necessary code to deal with these data can be found at: https://github.com/IvanMuro/agn_frac_data_release</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.