Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,326
datasets available to search
ShareScore release 0.9.0
Dataset results
2,326 results for “clusters”
Molecular Models and Wave Function Definitions for Models A-G of the [2Fe]F Cluster in FeFe-hydrogenase Maturase Enzyme HydF
<p>The dataset contains all relevant atomic positional coordinates for 2Fe-cluster models, and electronic wave function data (using formatted Gaussian checkpoint files) as described in the related publication (see citation below).</p> <p>The version 2.0 contains additional models for [2Fe-2S] cluster linked [2Fe]F constructs.</p> <p>The top folder contains "analysis.xlsx" electronic spreadsheet that summarizes all the numerical results for absolute and relative electronic energy values, internal coordinates, calculated and scaled vibrational frequencies for diatomic stretching modes. The details of developing scaled quantum forcefields as a function of level of theory and model composition are also given.<br> The schematic structural definitions are given in the "models.pdf" file and keys for abbreviations are provided in "symbols.txt" file.<br> </p>
Science ready spectra of star clusters and their best-fitting models described in the research paper "Using Star Clusters as Tracers of Star Formation and Chemical Evolution: the Chemical Enrichment History of the Large Magellanic Cloud" by Chilingarian & Asa'd
<p>Science ready spectra of star clusters in the Large Magellanic Cloud and their best-fitting templates (alpha-enhanced MILES based simple stellar population models) obtained using the NBursts full spectrum fitting code. Each spectrum is presented as a binary FITS table, which contains a spectrum (wavelength, flux, uncertainties), best-fitting template, best-fitting parameters (radial velocity, age, metallicity), and a pixel mask used in the fitting procedure. For each cluster, 5 spectra are provided, which correspond to [alpha/Fe] values from 0.0 to 0.4 dex with a step of 0.1 dex. The only exception is NGC2249, for which only 3 models are provided. The alpha-enhancement value of a model grid used in the fitting procedure is given in the FITS keyword MGFEGRID.</p>
Research institutions clustering based on the intensity of academic collaboration
<p>The clustering of research institutions has been conducted using the Louvain modularity algorithm. The Louvain modularity is a state-of-the-art method of identifying communities (clusters) in large networks. Modularity is a value between -1 and 1 that measures the density of edges inside communities to edges outside communities. Optimizing this value results in the best possible grouping of the nodes of a given network.</p> <p>In our exercise, the Louvain methods were applied to identify clusters of institutions within ACM and SSRN networks. In the network, nodes are constituted of institutions, and edges are represented by the intensity of research collaboration measured by number of papers co-authored by authors affiliated with the institutions.</p> <p>As an example, including the paper: <em>Fast unfolding of communities in large networks</em>, written by V. D. Blondel (Universite Catholique de Louvain), J-L. Guillaume (Universite Pierre et Marie Curie), R. Lambiotte (Imperial College London) and Etienne Lefebvre (Universite Catholique de Louvain) would impact the number of edges in our analysis in the following way:</p> <p>“Universite catholique de Louvain” ⇔ “Imperial College London” =+1</p> <p>“Universite catholique de Louvain” ⇔ “Universite Pierre et Marie Curie” =+1</p> <p>“Imperial College London” ⇔ “Universite Pierre et Marie Curie” =+1</p> <p>In our largest network we analyse 5362 institution nodes with 147 482 edges. The number of identified clusters highly depends on the resolution parameter. Resolution is a parameter for the Louvain community detection algorithm that affects the size of the recovered clusters. Smaller resolutions recover smaller, and therefore a larger number of clusters, and conversely, larger values recover clusters containing more data points. In all clusterizations, we have used a default resolution (1.0) tuned in the popular Gephi software for network analysis. Resolutions equal to one result in a moderate number of clusters, characterised by satisfactory statistical distribution. </p> <p><strong>Source:</strong></p> <p>- Association for Computing Machinery (ACM)</p> <p>Characteristics of the ACM Data Set following geographical classification</p> <p>Number of institutions: 5477</p> <p>Number of papers: 674684</p> <p>Number of countries: 122</p> <p>Years: 2011-2018</p> <p>As ACM contains publications across various areas of computer science, a more in-depth analysis requires the classification of papers into fields of interests. During the analysis, we looked at 3 wide areas:</p> <ul> <li> <p>Artificial intelligence and machine learning</p> </li> <li> <p>Technology (hardware, emerging technologies, infrastructure)</p> </li> <li> <p>Social issues</p> </li> </ul> <p>The 3 categories were set following expert analysis of the 1000 most frequent keywords in the dataset. If a term from the following list appeared among the paper’s keywords, the paper was assigned to that group, allowing a paper to assign to more than one group.</p> <p><strong>Files:</strong></p> <p>mod_ai.csv (based on keywords related to artificial intelligence)</p> <p>mod_tech.csv (based on keywords related to technologies)</p> <p>mod_soc.csv (based on keywords related to social issues)</p> <p>mod_all.csv (based on all papers)</p> <p> </p>
Datasets for Watset: Local-Global Graph Clustering with Applications in Sense and Frame Induction
<p>This dataset supplements the article “<a href="https://doi.org/10.1162/COLI_a_00354">Watset: Local-Global Graph Clustering with Applications in Sense and Frame Induction</a>” published in the Computational Linguistics journal:</p> <ul> <li> <p><code>watset-coli-lcc-performance.tsv</code>: runtime analysis</p> </li> <li> <p><code>watset-coli-synsets.zip</code>: synset induction experiment (note that <code>pairwise-{en-babelnet,ru-rwn}.pkl</code> files are excluded due to the licensing issues)</p> </li> <li> <p><code>watset-coli-triframes.zip</code>: semantic frame induction experiment</p> </li> <li> <p><code>watset-coli-classes.zip</code>: semantic class induction experiment</p> </li> </ul>
Simulated dataset (Almeida et al., 2018 - GigaScience) treated with various clustering programs to evaluate ReClustOR efficiency and constitency
<p>ReClustOR is a novel clustering method that overcomes some of the problems associated with classical ‘<em>heuristic’</em>clustering methods and consequently increases the stability and quality of the reconstructed OTUs. Moreover, the OTUs database defined with ReClustOR can be used as reference(s) with gradual enrichment of it, with new studies and samples. In this way, huge datasets like<em> </em>the Earth Microbiome Project can be easily used as references for smaller projects, thereby increasing the quality of comparisons between studies and datasets</p> <p>Here, we propose a new approach called ReClustOR (for RE-CLUSTering method using an Open-Reference approach) to improve OTU consistency (see https://doi.org/10.5281/zenodo.2597402). This new strategy combines two of the previously-described clustering methods. Firstly, a classical clustering<em> </em>method (<em>e.g. </em>SWARM, or VSEARCH) is used to define OTU centroids and create a reference database. Secondly, a closed- or open-reference method (depending on the user’s choice) is computed for all reads which are not considered as OTU centroids. Contrary to the classical clustering methods, each read is compared to all centroids using a distance-based greedy clustering technique (Edgar, 2010; He et al., 2015), and then assigned to the nearest one, thereby fixing the erroneous assignments of reads to OTUs.</p> <p>To highlight the improvements provided by ReClustOR in describing microbial diversity in terms of ecological diversity metrics (<em>e.g. </em>richness, OTU composition, Shannon, 1/Simpson) and taxonomic composition, a simulated dataset was subjected to: (i) ESV definition, (ii) multiple conventional <em>de novo</em> methods (<em>i.e.</em> a homemade <em>de novo </em>clustering close to CRUNCHCLUST, VSEARCH and SWARM), and (iii) ReClustOR computation. This dataset is a simulated one (Almeida et al., 2018), containing a diverse set of <em>genera</em> commonly found in three ecosystems different ecosystems: human gut, ocean and soil. The clustering methods were compared for: (i) their ability to describe microbial richness, (ii) the congruence between OTU assignments and sequences taxonomy, (iii) the robustness of each defined OTU, and (iv) their ability to efficiently describe the microbial community based on OTU composition.</p> <p>Here, the simulated dataset (00_Raw_data) and all steps of analysis are available to resue them to test ReClustOR, and also to have a better understanding of files and data produced by this program. More details are available in the Tree_of_data.tree file.</p>
Science ready spectra and their best-fitting models described in the research paper ``Internal dynamics and stellar content of nine ultra-diffuse galaxies in the Coma cluster prove their evolutionary link with dwarf early-type galaxies'' by Chilingarian et al.
<p>Science ready spectra of nine ultra-diffuse galaxies in the Coma cluster collected with the Binospec multi-object spectrograph and their best-fitting PEGASE.HR templates obtained using the NBursts full spectrum fitting code. These spectra were presented in the paper ``Internal dynamics and stellar content of nine ultra-diffuse galaxies in the Coma cluster prove their evolutionary link with dwarf early-type galaxies'' by Chilingarian et al. accepted for publication in the Astrophysical Journal on Sep/3/2019 (arXiv:1901.05489).</p> <p>Each spectrum is presented as a binary FITS table, which contains a spectrum (wavelength, flux, uncertainties), best-fitting template, best-fitting parameters (radial velocity, age, metallicity), and a pixel mask used in the fitting procedure. For six galaxies there are two files provided: (i) one-dimensional optimally extracted integrated spectrum and (ii) two dimensional spectrum for spatially resolved radial velocity information. For the remaining three galaxies, only spatially resolved spectra are provided.</p>
The Structure of Sub-nm Platinum Clusters at Elevated Temperatures (Supplementary Information)
<p><strong><em>This dataset consists of raw data and denoised scanning transmission electron microscopy videos of sub-nm sized clusters of Pt on a carbon substrate. The data is used in the article "The Structure of Sub-nm Platinum Clusters at Elevated Temperatures" published in Angewandte Chemie International Edition, 2019, DOI:10.1002/anie.201911068</em></strong><strong><em> </em></strong></p> <p><strong>Video S1.</strong> A typical sub-nm amorphous cluster at room temperature. 0.5 nm scale bar.</p> <p><strong>Video S2.</strong> Two typical crystalline sub-nm clusters at 350°C. 0.5 nm scale bar.</p> <p><strong>Video S3. </strong>In this high-speed recording at 147 fps, the high beam current required for this fast imaging has suppressed the crystallinity of the cluster, despite the temperature of 350°C. 0.5 nm scale bar.</p> <p><strong>Video S4. </strong>The unusually stable 13-atom cluster at the bottom forms an fcc cuboid, and can be seen rotating at three orientations, as shown by the inset model and in Fig. 3a-c. 0.5 nm scale bar.</p> <p><strong>Video S5. </strong>This 15-atom cluster initially forms an fcc cube, then transforms into multiple hcp structures. (Recorded at 2 fps, but animated at 5x real time at 10fps). 0.5 nm scale bar.</p> <p><strong>Video S6. </strong>The cluster in this movie is a 22-atom truncated rectangular cuboid. 0.5 nm scale bar.</p> <p><strong>Video S7. </strong>In the center and bottom, two 6-atom octagons are rotating (shown in Fig. S4) as they add onto their larger neighboring clusters. The 13-atom cluster in the top forms an unusually stable fcc cuboctahedron from frame 219. 0.5 nm scale bar.</p> <p><strong>Video S8. </strong>The cluster on the bottom left forms a fleeting icosahedron-like structure. 0.5 nm scale bar.</p> <p><strong>Video S9. </strong>This cluster shows fcc structures, despite being recorded at 200°C, but with a very low beam dose. (Recorded at 2 fps, but animated at 5x real time at 10fps). 0.5 nm scale bar.</p>
Empirical measurements of function placements and executions in a mixed cloud-edge cluster
<p>Empirical measurements used for the Skippy Scheduler, an optimized container scheduler for serverless edge computing in Kubernetes.</p>
Dataset and models of TMLR 2024 Paper "Identifying and Clustering Counter Relationships of Team Compositions in PvP Games for Efficient Balance Analysis"
<p>This is a part of dataset and models of the paper published in TMLR 2024 (Transactions on Machine Learning Research, <a href="https://jmlr.org/tmlr/" target="_blank" rel="noopener">https://jmlr.org/tmlr/</a>).</p> <p>Including training datasets, testing datasets, and models.</p> <p>The example program for using this file will be put on the author's github repo branch: <a href="https://github.com/DSobscure/cgi_drl_platform/tree/game_balance_measures_tmlr" target="_blank" rel="noopener">https://github.com/DSobscure/cgi_drl_platform/tree/game_balance_measures_tmlr</a></p> <p> </p>
The North Pacific Eukaryotic Gene Catalog: clustered nucleotide metatranscripts and read counts
<p>This data continues with the development of the NPEGC Trinity <em>de novo</em> metatranscriptome assemblies from the protein data repository of <a href="../doi/10.5281/zenodo.10472589">The North Pacific Eukaryotic Gene Catalog</a>. The nucleotide sequences corresponding to the NPEGC cluster representatives are collected together in these repository files:<br><br><em>NPac.G1PA.bf100.id99.nt.fasta.gz</em><br><em>NPac.G2PA.bf100.id99.nt.fasta.gz</em><br><em>NPac.G3PA.bf100.id99.nt.fasta.gz</em><br><em>NPac.G3PA_diel.bf100.id99.nt.fasta.gz</em><br><em>NPac.D1PA.bf100.id99.nt.fasta.gz</em><br><br>A full description of this data is published in Scientific Data, available here: <a href="https://www.nature.com/articles/s41597-024-04005-5" target="_blank" rel="noopener">The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations</a>. Please cite this publication if your research uses this data:<br><br>Groussman, R. D., Coesel, S. N., Durham, B. P., Schatz, M. J., & Armbrust, E. V. (2024). The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. <em>Scientific Data</em>, <em>11</em>(1), 1161.<br><br>These nucleotide sequences have been sourced from the Zenodo repository for raw assemblies: <a href="../records/7332796">The North Pacific Eukaryotic Gene Catalog: Raw assemblies from Gradients 1, 2 and 3</a></p> <p>Key processing steps are sampled below with links to the detailed code on the main github code repository: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog">https://github.com/armbrustlab/NPac_euk_gene_catalog</a></p> <p><br>Code used to build the kallisto indices and map the short reads against indices with kallisto are online in the code repository here: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/nt_data/NPEGC.nt_kallisto_counts.sh">NPEGC.nt_kallisto_counts.sh</a><br><br>There are two main steps:<br>1. Generate the kallisto index on the sets of clustered nucleotide metatranscripts<br>2. Map the short reads from environmental samples back to the assembly index</p> <p>As generated above, kallisto generates separate results files for each of the sample files. Even after compression, the total size of the tarballed kallisto output results directories are prohibitively large (>50GB). We use the code in this template R script to join together the 'est_count' estimated count values for the tens of millions of protein sequences in each project metatranscriptome, along with length.</p> <p>The code in this template script was used for each project: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/nt_data/aggregate_kallisto_counts.R">aggregate_kallisto_counts.R</a><br>The output count files for each project are Gzip-compressed and uploaded to the NPEGC nucleotide data repository here: </p> <p><em>G1PA.raw.est_counts.csv.gz</em><br><em>G2PA.raw.est_counts.csv.gz</em><br><em>G3PA.raw.est_counts.csv.gz</em><br><em>G3PA_diel.raw.est_counts.csv.gz</em><br><em>D1PA.raw.est_counts.csv.gz</em></p>
Global Cluster Test Results: Pathogenic Fungi in Decayed Norway Spruce Stands
<p><strong>Accessing the Results:</strong> Users can retrieve the test results by opening the dataset <code>Global_cluster_test_results.RData</code> in an R session and using the functions inside the script of the same name.</p> <p><strong>Description: </strong>This dataset provides global cluster test results analyzing the spatial distribution of pathogenic fungi in 273 Norway spruce stands in Norway (Lara et al., 2024). The stands, composed mainly of Norway spruce (27% to 100%), also include Scots pine and birch. It focuses on spatial patterns of decayed spruce trees, offering p-values, clustering metrics, and other parameters from statistical analyses.</p> <p><strong>Analysis:</strong> The dataset includes results from three global cluster tests (Tango, 2010):</p> <ul> <li>Tango's Nearest Neighbors (TNN)</li> <li>Tango's Double Exponential Clinal (TCN)</li> <li>Diggle and Chetwynd’s (DC)</li> </ul> <p><strong>Methodology:</strong> Cluster testing employed 1,000 Monte Carlo simulations for each test across all stands to establish null distributions and adjusted p-values, ensuring robust statistical assessments under the random labeling hypothesis: H0: the observed n0 decayed trees are a random sample from the entire sample of size n = n0 + n1 (decayed trees + healthy trees) (Tango, 2010).</p> <p> </p>
CLDF dataset derived from Othaniel's "Jen Cluster Comparative Wordlist" from 2017
<p>Cite the source of the dataset as:</p> <blockquote> <p>Othaniel, Nlabephee Kefas. 2017. A phonological comparative study of the Jen language cluster. (MA thesis, Jos: Theological College of Northern Nigeria; 1–83pp.)</p> </blockquote>
Reproduction package for the paper "Two waves of massive stars running away from the young cluster R136"
<h2>Reproduction package for the paper "Two waves of massive stars running away from the young cluster R136".</h2> <ul> <li>This reproduction package aims for open science, with the internal API designation of 'Gold'</li> <li>Authors: M. Stoop, A. de Koter, L. Kaper, S. Brands, S. Portegies Zwart, H. Sana, F. Stoppa, M. Gieles, L. Mahy, T. Shenar, D. Guo, G. Nelemans, S. Rieder</li> <li>Paper DOI: https://doi.org/10.1038/s41586-024-08013-8</li> <li>Zenodo DOI: http://doi.org/10.5281/zenodo.10058762</li> <li>Published in Nature (date of publication: 2024/10/09)</li> </ul> <h2>Hardware</h2> <ul> <li>Tested on a MacBook Pro (13-inch, 2020, Four Thunderbolt 3 ports)</li> <li>Processor: 2 GHz Quad-Core Intel Core i5</li> <li>Memory: 32 GB 3733 MHz LPDDR4X</li> <li>Graphics: Intel Iris Plus Graphics 1536 MB</li> </ul> <h2>Required non-standard hardware</h2> <ul> <li>None</li> </ul> <h2>Software dependencies</h2> <ul> <li>Jupyterlab (4.0.8)</li> <li>Notebook (7.0.6)</li> <li>Programming languages used: Python (3.11.7)</li> <li>Python packages used: numpy (1.25.2), pandas (2.1.4), matplotlib (3.8.0), os (comes with Python) scipy (1.11.4), gaiadr3-zeropoint (0.0.4) https://gitlab.com/icc-ub/public/gaiadr3_zeropoint), astroquery (0.4.6), pymc (5.6.1), corner (2.2.2), arviz (0.16.0), pytensor (2.12.3), lmfit (1.2.2), powerlaw (1.5), rpy2 (3.5.16), seaborn (0.12.2), consistencytest (0.0.2)</li> </ul> <h2>Instructions</h2> <ul> <li>The Anaconda conda environment is given should this be needed</li> <li>All Jupyter Notebooks are ready-made to produce the raw data, intermediate and end data products</li> <li>Gaia raw data is downloaded in the Jupyter Notebook "R136_runaway_candidates.ipynb"</li> <li>Data from the literature is given in the subdirectory /tables/ or /input_files/</li> <li>Input images and files are given in the subdirectory /input_files/</li> <li>Intermediate and end data products are given in /output_files/</li> <li>Figures in the paper are produced in the Jupyter Notebooks in the subdirectory /figures/ and stored in the subdirectory /figures/figures_paper/</li> </ul> <h2>Expected Output</h2> <ul> <li>Jupyter Notebooks can be executed by "Run" -> "Run All Cells"</li> <li>The Jupyter Notebook show the expected output in their respective cell</li> <li>Expected runtime are given at the top of each Jupyter Notebook</li> <li>The Jupyter Notebook which takes the longest "R136_runaway_search.ipynb" takes 7-8 hours for the entire dataset</li> <li>A small dataset has been given in this Jupyter Notebook as a proof-of-concept</li> </ul> <h2>Instructions for use</h2> <ul> <li>Jupyter Notebooks can be executed by "Run" -> "Run All Cells"</li> </ul> <h2>Figures</h2> <ul> <li>Figures can be reproduced from the /figures/ folder.</li> <li>All material and data used are available either in the Raw Data or in the Intermediate Data</li> <li>The figures shown in the paper will be saved in ./figures/figures_paper/ folder.</li> </ul> <pre> </pre>
Supplementary Material for "A comparative high-resolution spectroscopic analysis of in situ and accreted globular clusters"
<p>This is a file containing supplementary material for the paper <em>A comparative high-resolution spectroscopic analysis of in situ and accreted globular clusters.</em> For each star in target globular clusters, it lists crucial information on the linelist analyzed. In particular:</p> <ol> <li>Star ID.</li> <li>Chemical element.</li> <li>Wavelength.</li> <li>log <em>gf</em></li> <li>Excitation potential.</li> <li>Measured equivalent width with uncertaintiy.</li> </ol>
Magnetic arch plasma expansion in a cluster of two ECR plasma sources (RPA and FC measurements)
<p>- Data from: Magnetic arch plasma expansion in a cluster of two ECR plasma sources (RPA and FC measurements)</p> <p>- Authors: Célian Boyé, Jaume Navarro-Cavallé, Mario Merino</p> <p>- Contact email: <a href="mailto:cboye@ing.uc3m.es" target="_blank" rel="noopener">cboye@ing.uc3m.es</a></p> <p>- Date: 2024-10-24</p> <p>- Version: 1.0</p> <p>- License: This dataset is made available under the <a href="https://creativecommons.org/licenses/by/4.0/legalcode">Creative Commons Attribution 4.0 International</a></p> <p> </p> <h2>Abstract</h2> <p>This dataset contains the raw experimental data employed in:</p> <p>Célian Boyé, Jaume Navarro-Cavallé, Mario Merino, "Magnetic arch plasma expansion in a cluster of two ECR plasma sources", Journal of Electric Propulsion.</p> <p>Which is currently submitted.</p> <p> </p> <h2>Dataset description</h2> <p>The experimental data is gathered by means of a Retarding Potential Analyzer (RPA) and a Faraday Cup (FC). The probes have been set on a polar probing arm system to scan the central horizontal plane of the setup, aligned with the axis of symmetry of the assembly and pointing toward the origin at the exit plane of the source(s).</p> <p>The RPA data is provided separately for every spatial position inspected for each configuration (S0, S1, D0, DA, DB). It is collected by means of an Impedance-Semion Retarded Potential Analyser, with a mean resolving voltage of 1V. The FC data is provided for the DA configuration to support the RPA measurements. </p> <p>Please refer to the corresponding article for further details regarding the data collection.</p> <p> </p> <h2>Data files</h2> <p>The data files are in standard comma separated values .csv format. Many programming languages provide functionalities to load such fields.</p> <ul> <li> <h3>RPA data</h3> </li> </ul> <p>The RPA data is separated through the different configurations:</p> <ul> <li> <ul> <li>S0: single ECR source without applied magnetic field.</li> <li>S1: single ECR source with applied magnetic field.</li> <li>D0: cluster of ECR sources without applied magnetic field.</li> <li>DA: cluster of ECR sources with opposed polarity.</li> <li>DB: cluster of ECR sources with same polarity.</li> </ul> </li> </ul> <p>The angle steps vary through the different configurations. Each file contains 8 headlines. </p> <ul> <li> <ul> <li>The first column contains the voltage applied to the sweeping grid (V).</li> <li>The second to sixth columns contain the current collected by the collector (A).</li> <li>The eventh to eleventh columns contain the derivative of the collected current by the voltage (A/V).</li> </ul> </li> </ul> <ul> <li> <h3>FC data</h3> </li> </ul> <p>The FC data has been probed for the DA configuration. The file contains 2 headlines.</p> <ul> <li> <ul> <li>The first column contains the angle at which the current has been collected (deg).</li> <li>The second column contains the distance from the origin at the exit plane of the cluster (mm).</li> <li>The third column contains the collected current (A).</li> </ul> </li> </ul> <p> </p> <h2>Citation</h2> <p>Works using this dataset or any part of it in any form shall cite it as follows.</p> <p>The preferred means of citation is to reference the publication associated to this dataset, as soon as it is available.</p> <p>Optionally, the dataset may be cited directly by referencing the corresponding DOI: 10.5281/zenodo.13987138</p> <p> </p> <h2>Acknowledgments</h2> <p>This work has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (project ERC-STG ZARATHUSTRA, grant agreement No 950466). </p>
Twin Test 2: Wake interactions of a cluster of turbines and wake steering techniques. Wind tunnel data.
<p>The aerodynamic performance of two identical wind turbine models was characterized under various static and dynamic conditions in a synchronous configuration within the wind tunnel test section. Two experimental campaigns were performed at Technische Universität München (TUM) and at the National Technical University of Athens (NTUA) to investigate wake flow control techniques. This document contains the necessary information to understand the performed experiments and to access and use the available data. While both experimental set ups are detailed, only data from the TUM campaign are available at the time of writing, as the NTUA campaign results will form Phase II of an ongoing blind test campaign and cannot be published.</p>
Spatial clustering of Neobuccinum eatoni occurrence data for potential distribution modeling
<p>The occurrence dataset for <em>Neobuccinum eatoni</em> was compiled through filtration process, starting with records from the Global Biodiversity Information Facility (GBIF) and supplemented by museum specimens and additional sources like SOMBASE, iBOL, NIWA, ANTABIF, and SCAR-AntOBIS. Further data were sourced from the National Museum of Natural History in Paris, the University of Vigo, and recent fieldwork in Antarctica, Heard Island, and Kerguelen Island. Records were meticulously screened to remove misidentified specimens, inaccurate locations, duplicates, and outdated entries, ensuring accuracy and relevance. To address spatial autocorrelation, clustering methods divided the data into distinct geographic clusters, producing a refined dataset used to model <em>N. eatoni</em>'s potential distribution with enhanced predictive reliability by reducing spatial autocorrelation effects.</p>
Artifact Description/Artifact Evaluation/Computational Artifact for paper, entitled Analytic Roofline Modeling and Energy Analysis of the LULESH Proxy Application on Multi-Core Clusters
We provide reproducibility initiative dependencies (Artifact Description or Artifact Evaluation or Computational Results Analysis) appendix. To allow a third party to duplicate the findings, this article provides our extensive performance data artifact and describes further details regarding the software environments, experimental design, and methodology employed for the results shown in the paper. The computational artifacts will enable experienced performance engineers to reproduce and interpret the data shown in the paper in the appropriate way and to follow the conclusions we draw from it.
Codes and datasets for a brief introduction to clustering and dimensionality reduction
<p>Python codes and datasets for the examples on unsupervised learning presented in Joris Paret's doctoral thesis « Hidden order in disordered materials » (2021).</p>
A computational intelligence approach to predict energy demand using Random Forest in a Cloudera cluster
<p>Society’s energy consumption has shot up in recent years, making the prediction of its demand a current challenge to ensure an efficient and responsible use. Artificial intelligence techniques have proven to be potential tools in handling tedious tasks and making sense of large-scale data to make better business decisions in different areas of knowledge. In this article, the use of random forests algorithms in a Big Data environment is proposed for households energy demand forecasting. The predictions are based on the use of information from different sources, confirming a fundamental role of socioeconomic data in consumer’s behaviours. On the other hand, the use of Big Data architectures is proposed to perform horizontal and vertical scaling of the solution to be used in real environments. Finally, a tool for high-resolution predictions with great efficiency is introduced, which enables energy management in a very accurate way.</p> <p>Raw data is incuded in data.csv. This file contains half hourly home electricity consumption registers for 4404 households with fix tariffs (not subject to dynamic time of use) for a period between November 2011 and February 2014. Original information was acquired from the Low Carbon London project led by UK Power Networks (https://data.london.gov.uk/dataset/smartmeter-energy-use-data-in-london-households)</p> <p>RFResults.zip contains the energy predictions for each ACORN group using the generated Random Forest algorithm. For this purpose, the first 613 days of a total of 818 observations of each group were considered for training and the last 205 days for testing.</p> <p>Meteorological data was adquired from the darksky app (https://darksky.net). These data are included in the weather_hourly_darksky.csv</p> <p>uk_bank_holidays. xlsx contains the dated of UK bank holidays for the studied period, used as additional variable related to occupancy</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.