Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11
datasets available to search
ShareScore release 0.9.0
Dataset results
11 results for “task cluster”
Task resource usage of Google Cluster Usage Trace dataset
<p>The dataset contains csv files. All csv files named by its cluster number and machine id. There are 2 features , time stamp and mean CPU usage. This data set is prepared from task resource usage (TRU) table of GCT dataset for the research purpose. Mainly TRU contains task information. Sum Average algorithm employed and calculate the machine information for every 5 minute time stamp.</p>
Companion data of Summarizing task-based applications behavior over many nodes through progression clustering
<p>This is the companion data for the paper *Summarizing task-based applications behavior over many nodes through progression clustering* by Lucas Leandro Nesi, Vinícius Garcia Pinto, Lucas Mello Schnorr, and Arnaud Legrand accept for publication in 31st Euromicro International Conference on Parallel, Distributed, and Network-Based Processing (<a href="https://www.pdp2023.org/">PDP 2023</a>). The remaining of the companion is at: https://gitlab.com/lnesi/companion-pdp-2023.</p>
Companion data of Detection, Evaluation and Mitigation of Resource Affinity and Communication Contention Problems in a Task-Based Runtime over Heterogeneous Clusters
<p>This is the companion data repository for the paper entitled <strong>Detection, Evaluation, and Mitigation of Resource Affinity and Communication Contention Problems in a Task-Based Runtime over Heterogeneous Clusters</strong> by Lucas Leandro Nesi and Lucas Mello Schnorr. The manuscript has been accepted for publication in the <a href="http://wscad.sbc.org.br/current/index.html">WSCAD 2020</a>.</p>
Clustering Tasks and Decision Trees with Elegiac Poets
<p>The dataset contains files generated during a Natural Language Processing (NLP) and automatic text analysis task. Attached is a <strong>Jupyter notebook</strong> with the complete code, along with several <strong>Excel files (.xlsx)</strong> containing organized information. Additionally, there are three folders that include files generated during the Silhouette calculation, K-means clustering, and feature extraction using decision trees.</p> <p>The three folders are:<br>1. <strong>Silhouette Calculation:</strong> Contains PNG images of Silhouette plots for various analysis configurations.<br>2.<strong> K-means Clustering: </strong>Contains pickle (.pkl) files with features and labels for each combination of excluded author, n-gram type, n-gram range, and matrix type.<br>3. <strong>Feature Extraction: </strong>Contains CSV files with lists of documents by cluster and the most important features along with information gain and information gain ratio metrics.</p> <p>Other file formats included in the dataset are:<br>- CSV files containing Silhouette scores, optimal clustering results, cluster assignments, and optimal cluster assignments.<br>- PNG images of scatter plots colored by author and by cluster.<br>- Pickle files containing the top features extracted during the analysis.</p>
Data from: Health insurance coverage with or without a nurse-led task shifting strategy for hypertension control: a pragmatic cluster randomized trial in Ghana
Open the record for dataset details and reuse information.
Figure 4 from: Nelson G, Paul D, Riccardi G, Mast A (2012) Five task clusters that enable efficient and effective digitization of biological collections. ZooKeys 209: 19-45. https://doi.org/10.3897/zookeys.209.3135
Figure 4 - Specimen image processing. Using Adobe Photoshop Lightroom software to process images. New York Botanical Garden.
Figure 3 from: Nelson G, Paul D, Riccardi G, Mast A (2012) Five task clusters that enable efficient and effective digitization of biological collections. ZooKeys 209: 19-45. https://doi.org/10.3897/zookeys.209.3135
Figure 3 - Custom specimen holder. Museum of Compartive Zoology (MCZ) Rhopalocera (Lepidoptera) Rapid Digitization Project.
Figure 2 from: Nelson G, Paul D, Riccardi G, Mast A (2012) Five task clusters that enable efficient and effective digitization of biological collections. ZooKeys 209: 19-45. https://doi.org/10.3897/zookeys.209.3135
Figure 2 - Specimen image capture. Fossil specimen imaging, specimen label imaging. Two very different imaging set-ups. Yale Peabody Museum, University of Kansas - Entomology.
Figure 1 from: Nelson G, Paul D, Riccardi G, Mast A (2012) Five task clusters that enable efficient and effective digitization of biological collections. ZooKeys 209: 19-45. https://doi.org/10.3897/zookeys.209.3135
Figure 1 - Pre-digitization specimen curation and staging. Preparing barcodes and imaging labels, affixing barcodes, updating taxonomy. L to R: University of Kansas – Entomology, New York Botanical Garden and Yale Peabody Museum.
Figure 5 from: Nelson G, Paul D, Riccardi G, Mast A (2012) Five task clusters that enable efficient and effective digitization of biological collections. ZooKeys 209: 19-45. https://doi.org/10.3897/zookeys.209.3135
Figure 5 - Electronic data capture. Entering data straight from the specimen label into the database. New York Botanical Garden.
Figure 6 from: Nelson G, Paul D, Riccardi G, Mast A (2012) Five task clusters that enable efficient and effective digitization of biological collections. ZooKeys 209: 19-45. https://doi.org/10.3897/zookeys.209.3135
Figure 6 - Dominant Digitization Workflows Observed.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.