Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,326

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,326 results for “clusters”

Learn how ShareScore rates datasets ↗
zenodo32/100

Clustering Dataset

<pre>This dataset contains jointly optimised energy system model networks on transmission and generation level with the objective to minimise total system costs under the constraint to reduce 95% of CO2 emissions compared to 1990 by a greenfield approach. We vary the spatial resolution on different scales (clustering on generation sites, the transmission network and joint clustering) to research the impact of each scale separately for the optimisation process. For each resolution, we additionally consider seven different transmission line expansion volumes. Results indicate, that&nbsp;in the current European transmission system, transmission network bottlenecks play a more restricting role than resource availability. Accounting for bottlenecks can raise total system costs by up to 19%, while a higher resource availability (locations with higher capacity factors) mitigates the increase only by 6.5%. Both cases result in strong differences in the composition of technology investments. A detailed survey of the results will be presented in the paper &quot;The strong effect of network resolution on electricity system models with high shares of wind and solar&quot;.</pre>

opencc-by-4.0Jul 2020View details →
zenodo32/100

Reproduction package for "X-ray study of the merging galaxy cluster Abell 3411-3412 with XMM-Newton and Suzaku"

<p>This is the reproduction package of the paper&nbsp;&quot;X-ray study of the merging galaxy cluster Abell 3411-3412 with XMM-Newton &nbsp;and Suzaku&quot; (arXiv:2007.15976).&nbsp;</p>

opencc-by-4.0Aug 2020View details →
dryad32/100

Divalent cations bind to phosphoinositides to induce ion and isomer specific propensities for nano-cluster initiation in bilayer membranes

<p>We report all-atom molecular dynamics simulations of physiologically composed asymmetric bilayers containing phosphoinositides in the presence of monovalent and divalent cations. We have characterized the molecular mechanism by which these divalent cations interact with phosphoinositides. Calcium desolvates more readily, consistent with single-molecule calculations, and forms a network of ionic-like bonds that serve as a "molecular glue'' that allows a single ion to coordinate with up to three phosphatidylinositol-(4,5)-bisphosphate lipids. The phosphatidylinositol-(3,5)-bisphosphate isomer shows no such effect and neither does PI(4,5)P2 in the presence of Mg. The resulting network of Ca-mediated lipid-lipid bonds grows to span the entire simulation space and therefore has implications for the lateral distribution of phosophoinositides in the bilayer. We observe context-specific differences in lipid diffusion rates, lipid surface densities, and bilayer structure. The molecular-scale delineation of ion-lipid arrangements reported here provides insight into similar nanocluster formation induced by peripheral proteins to regulate the formation of functional signaling complexes on the membrane.</p>

opencc-zeroAug 2020View details →
zenodo32/100

Supervised jet clustering reference data

<p>A set of MC simulated events for training/testing machine learning architectures for the task of supervised jet clustering.<br> In total 100,000 W&#39; events and 100,000 q* events<br> Description:<br> * 13 TeV collision data simulated with pythia 8.183.<br> * wboson.txt contains events generated from a W&#39; boson with a mass of 600 GeV, which decays 100% of the time to a W boson and a Z boson. The W boson is forced to decay haronically and the Z boson decays into neutrinos.<br> * qstar.txt contains events generated from a excited quark q* with a mass of 600 GeV, which decays 100% of the time to a quark and a Z boson. The Z boson is forced to decay into neutrinos.<br> * events in the text format<br> * each line in the text represent one event, which contains variable number of detector-stable particles.<br> * each particle contains 7 features in order: [px, py, pz, E, pdgID, is-from-W, is-in-leading-jet]. The first four features are the four momentum of the particle, and pdgID is the pag number of the particle. is-from-W is 1 if the particle coming from W boson and 0 otherwise. is-in-leading-jet is 1 if the particle is inside the leading jet reconstructed from the anti-kT jet algorithm (R=1.0).</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2020View details →
zenodo32/100

Audiovisual integration of single consonants and consonant clusters (Audiovisual dataset)

<p>Audiovisual speech stimuli generated for the manuscript &ldquo;Audiovisual integration of single consonants and consonant clusters.&rdquo;</p> <p>The file &ldquo;AV_stimuli.zip&rdquo; contains congruent and incongruent (McGurk) audiovisual speech stimuli, which were generated with all possible audiovisual pairings of the disyllables /aba/, /aga/, /ada/, /abga/ and /abda/. The utterances were produced by a female native speaker of Swedish. The video recordings were made at a resolution of 1920 x 1080 and 30 frames per second. The sound was recorded at a bit depth of 24 bits and a sampling rate of 48 kHz.&nbsp;</p>

opencc-by-4.0Aug 2020View details →
zenodo32/100

Cluster der Indien-Topik mit einer Auswahl auffälliger Kollokationsmuster (Abb. 2)

<p>Diese Abbildung ist Teil des folgenden Buchs: <a href="https://www.transcript-verlag.de/978-3-8376-5227-7/topik-zwischen-modellierung-und-operationalisierung/">https://www.transcript-verlag.de/978-3-8376-5227-7/topik-zwischen-modellierung-und-operationalisierung/</a>. Die Darstellung ist inhaltlich identisch mit Abb. 2 des verlinkten Buchs, wird jedoch hier als unzerteilte Visualisierung in einer Querformatansicht zug&auml;nglich gemacht.</p> <p>Die Studie analysiert Topoi als dynamische Kristallisationspunkte, &uuml;ber die sich argumentative Muster begreifen lassen. Es wird ein Ansatz verfolgt, der die Relationalit&auml;t und Kombinatorik zentral setzt: Topoi sind nicht isoliert wirksam, sondern treten in konkreten sprachlichen Zusammenh&auml;ngen kombiniert auf; sie sind in <em>Topiken</em> verankert, die mittels der Topik als Heuristik untersucht werden. F&uuml;r das Untersuchungskorpus von rund 40 Indienreiseberichten um 1900 wurden in der (Re‑)Konstruktionsarbeit zwei <em>Topiken</em> als relevant erachtet: Die Indien-<em>Topik</em> setzt sich aus 86 Topoi zusammen, die zur topischen Stabilisierung eines imagin&auml;ren &sbquo;Indien&lsquo; um 1900 beitragen, wohingegen die Reiseberichts-<em>Topik</em> eine Konstellation von 67 Topoi bildet, die als charakteristisch f&uuml;r die Textsorte Reisebericht um 1900 gelten k&ouml;nnen. Um mit einer solchen Menge an Topoi analytisch-interpretativ arbeiten zu k&ouml;nnen, erfolgt in einem weiteren Schritt die Clusteranalyse, in deren Rahmen funktional homogene Topoi zu Clustern gruppiert werden.</p> <p>Diese Visualisierung bildet die 86 Topoi systematisch ab und veranschaulicht einen zentralen Befund: F&uuml;r die Indien-Topik sind zwei grundlegend verschiedene Clustertypen (&sbquo;Inventar&lsquo;-Cluster und &sbquo;thematisch-diskursive&lsquo; Cluster) sowie spezifische Muster topischer Kollokationen als zentrales Ergebnis hervorzuheben.</p> <p>Die Graphik wurde mit der Visualisierungs- und Analysesoftware VUE (= <em>Visual Understanding Environment</em>, ein Open Source Projekt der Tufts University, <a href="https://vue.tufts.edu/">https://vue.tufts.edu/</a>) erstellt.</p>

opencc-by-4.0Sep 2020View details →
zenodo32/100

First COVID-19 Genomic Patient Cluster was at PLA Hospital in Wuhan, China

<p>A paper published on Zenodo (DOI 10.5281/zenodo.4119263) by Dr. Steven Quay, M.D., PhD., head of two COVID-19 therapeutic programs at Atossa Therapeutics, Inc. (NASDAQ: ATOS), illuminates new scientific observations and conclusions documenting that the SARS-CoV-2 pandemic began at the General Hospital of Central Theater Command of People&rsquo;s Liberation Army (PLA Hospital) in Wuhan, China, located at 627 Wulon Road, Wuchang District, Wuhan. International biospecimen data repositories indicate as early as December 10, 2019 COVID patient records were being created by PLA personnel, weeks before the Chinese government informed the WHO of the pandemic.</p> <p>&nbsp;</p> <p>This is a short presentation by the author explaining his research.</p>

opencc-by-4.0Oct 2020View details →
zenodo32/100

Smart Energy Use Case Clustering Dataset - II

<p>Cluster data set for multilayered and contextualized smart energy blockchain traces</p>

opencc-by-4.0Nov 2020View details →
zenodo32/100

Cluster configurations of the Hegselmann-Krause model on network ensembles

<p>This is the raw data underlying the results of the preprint [arxiv:2102.10910](https://arxiv.org/abs/2102.10910).</p> <p>&nbsp;</p> <p>## Data</p> <p>For each measured combination of the confidence and system size, there is one gzipped<br> file. For different ensembles, we collected data in different ranges and quality.<br> The paramters are:</p> <p>* Number of samples `m` per parameter combination<br> * Range `r` of confidences epsilon<br> * Distances `d` between values of epsilon (basically the resolution of the data)<br> * Largest size `N_max`</p> <p>The single files follow a naming scheme of `n{N}_e{epsilon}.cluster.dat.gz`, where<br> `{N}` signals the system size of the simulation and `{epsilon}` is the confidence<br> value of the simulation (without a decimal point, i.e., `0050` corresponds to `epsilon = 0.050`).<br> The sizes `N` are usually powers of two (or for the lattices, perfect squares close to powers of two).</p> <p>We present the data for each ensemble in one archive.</p> <p><br> * Fully connected `full.tar`<br> &nbsp;&nbsp;&nbsp; * `m = 1000`, `r = [0.0, 0.6]`, `d = 0.001`, `N_max = 262144`<br> * Barabasi Albert with a mean degree of 4 `BA4.tar`<br> &nbsp;&nbsp;&nbsp; * `m = 1000`, `r = [0.0, 0.6]`, `d = 0.002`, `N_max = 32768`<br> * Barabasi Albert with a mean degree of 10 `BA10.tar`<br> &nbsp;&nbsp;&nbsp; * `m = 1000`, `r = [0.0, 0.6]`, `d = 0.001`, `N_max = 65536`<br> * Square lattice with first nearest neighbors `lat1.tar`<br> &nbsp;&nbsp;&nbsp; * `m = 1000`, `r = [0.0, 0.6]`, `d = 0.001`, `N_max = 16384`<br> * Square lattice with second nearest neighbors `lat2.tar`<br> &nbsp;&nbsp;&nbsp; * `m = 1000`, `r = [0.0, 0.6]`, `d = 0.001`, `N_max = 16384`<br> * Square lattice with third nearest neighbors `lat3.tar`<br> &nbsp;&nbsp;&nbsp; * `m = 1000`, `r = [0.0, 0.6]`, `d = 0.001`, `N_max = 65536`<br> * Square lattice with fourth nearest neighbors `lat4.tar`<br> &nbsp;&nbsp;&nbsp; * `m = 1000`, `r = [0.0, 0.6]`, `d = 0.001`, `N_max = 65536`<br> * Square lattice with third nearest neighbors and 1% rewired edges `lat3_ws.tar`<br> &nbsp;&nbsp;&nbsp; * `m = 1000`, `r = [0.0, 0.3]`, `d = 0.001`, `N_max = 16384`<br> * connected Erdos Renyi with mean degree of 10 `ER10.tar`<br> &nbsp;&nbsp;&nbsp; * `m = 1000`, `r = [0.0, 0.3]`, `d = 0.002`, `N_max = 32768`</p> <p>&nbsp;</p> <p>## Data format</p> <p>Each final state is encoded as three lines:</p> <p>* The convergence time is a single integer with a line prefix &#39;# sweeps: &#39;<br> * The positions of all clusters in opinion space with a line prefix &#39;# &#39; (unsorted)<br> * The number of agents in each of the clusters without a line prefix</p> <p>&nbsp;</p> <p>## Python example for reading the format</p> <p>An example script, which visualizes the S vs eps graph for the largest size of the fully connected<br> case, with a function to read this format is given in `example.py`.</p>

opencc-by-4.0Nov 2020View details →
zenodo32/100

Data and Scripts for "Real-Time Time-Dependent Density Functional Theory Implementation of Electronic Circular Dichroism Applied to Nanoscale Metal-Organic Clusters"

<p>The xyz files were used for the ECD calculation in the paper :https://arxiv.org/abs/2007.08560<br> gsrun.py: ground-state calculation<br> td_calc.py: time -propagation for ECD<br> cdspecX.py: get rotatory strength for plotting</p>

opencc-by-4.0Nov 2020View details →
zenodo32/100

Clustered Embedding using Deep Learning to Analyze Urban Mobility based on Complex Transportation Data

<p>The subset of the anonymized dataset for personalized POI embedding.</p>

opencc-by-4.0Dec 2020View details →
dryad32/100

Plasmodium falciparum genomic surveillance reveals spatial and temporal trends, association of genetic and physical distance, and household clustering

<p>Molecular epidemiology using genomic data can help identify relationships between malaria parasite population structure, malaria transmission intensity, and ultimately help generate actionable data to assess the effectiveness of malaria control strategies. Genomic data, coupled with geographic information systems data, can further identify clusters or hotspots of malaria transmission, parasite genetic and spatial connectivity, and parasite movement by human or mosquito mobility over time and space.  In this study, we performed longitudinal genomic surveillance in a cohort of 70 participants over four years from different neighborhoods and households in Thiès, Senegal—a region of exceptionally low malaria transmission (entomological inoculation rate (EIR) less than 1). Genetic identity (identity by state) was established using a 24 single nucleotide polymorphism molecular barcode and a multivariable linear regression model was used to establish genetic and spatial relationships. Our results show clustering of genetically similar parasites within households and a decline in genetic similarity of parasites with increasing distance.  One household showed extremely high diversity and warrants further investigation as to the source of these diverse genetic types. This study illustrates the utility of genomic data with traditional epidemiological approaches for surveillance and detection of trends and patterns in malaria transmission not only by neighborhood but also by household. This approach can be implemented regionally and countrywide to strengthen and support malaria control and elimination efforts.     </p>

opencc-zeroDec 2020View details →
zenodo32/100

Variable-order fractional master equation and clustering of particles: non-uniform lysosome distribution

<p>Raw and tracking data to accompany publication titled: &#39;Variable-order fractional master equation and clustering of particles: non-uniform lysosome distribution&#39;</p> <p>arXiv preprint available at&nbsp;https://arxiv.org/abs/2101.02698</p>

opencc-by-4.0Jan 2021View details →
zenodo32/100

The Altotiberina Low-angle normal fault and Gubbio fault in seismic cluster of 2014-2015 period.

<p>Seismology data fron INGV and ISC catalogue.</p>

opencc-by-4.0Oct 2017View details →
dryad32/100

Data from: Evidence for a stochastic geometry of biodiversity: the effects of species abundance, richness and intraspecific clustering

Most ecological theories that aim to explain coexistence in megadiverse communities employ a set of three rules to describe the stochastic geometry of biodiversity: (i) individuals exhibit intraspecific clustering; (ii) species abundances vary according to a log-normal distribution and (iii) the spatial arrangement between species is independent. The first two rules have received strong empirical support, but the third remains largely unexplored. To address this deficiency, we evaluated the independent species arrangement rule in a species-rich shrubland and its potential drivers, that is, the levels of species richness and intraspecific clustering exhibited by a given species at different scales, and the relative abundance of such species in the community. We found that interspecific associations were rare and that independence was positively related to species richness and intraspecific clustering, but negatively related to relative species abundances. Synthesis. Our results agree with the independent species arrangement rule and they provide empirical support for the stochastic geometry of biodiversity. In the context of species-rich plant communities, the likelihood of two species encountering is very small. However, our study demonstrated a novel feature of this context, where both intraspecific clustering (due limitations on dispersal) and relative species abundances play fundamental roles in determining the probability of two species encountering and interacting, especially at very fine spatial scales.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Population genomic analyses reveal a highly differentiated and endangered genetic cluster of northern goshawks (Accipiter gentilis laingi) in Haida Gwaii

Accurate knowledge of geographic ranges and genetic relationships among populations is important when managing a species or population of conservation concern. Along the western coast of Canada, a subspecies of the northern goshawk (Accipiter gentilis laingi) is legally designated as Threatened. The range and distinctness of this form, in comparison to the broadly distributed North American subspecies (Accipiter gentilis atricapillus), is unclear. Given this morphological uncertainty, we analyzed genomic relationships in thousands of single nucleotide polymorphisms identified using genotyping-by-sequencing of high-quality genetic samples. Results revealed a genetically distinct population of northern goshawks on the archipelago of Haida Gwaii and subtle structuring among other North American sampling regions. We then developed genotyping assays for ten loci that are highly differentiated between the two main genetic clusters, allowing inclusion of hundreds of low-quality samples and confirming that the distinct genetic cluster is restricted to Haida Gwaii. As the laingi form was originally described as being based in Haida Gwaii (where the type specimen is from), further morphological analysis may result in this name being restricted to the Haida Gwaii genetic cluster. Regardless of taxonomic treatment, the distinct Haida Gwaii genetic cluster along with the small and declining population size of the Haida Gwaii population suggests a high risk of extinction of an ecologically and genetically distinct form of northern goshawk. Outside of Haida Gwaii, sampling regions along the coast of BC and southeast Alaska (often considered regions inhabited by laingi) show some subtle differentiation from other North American regions. These results will increase the effectiveness of conservation management of northern goshawks in northwestern North America. More broadly, other conservation-related studies of genetic variation may benefit from the two-step approach we employed that first surveys genomic variation using high-quality samples and then genotypes low-quality samples at particularly informative loci.

opencc-zeroDec 2017View details →
dryad32/100

Data from: The impact of hotspot-targeted interventions on malaria transmission in Rachuonyo south district in the western Kenyan highlands: a cluster-randomized controlled trial

Background: Malaria transmission is highly heterogeneous, generating malaria hotspots that can fuel malaria transmission across a wider area. Targeting hotspots may represent an efficacious strategy for reducing malaria transmission. We determined the impact of interventions targeted to serologically defined malaria hotspots on malaria transmission both inside hotspots and in surrounding communities. Methods and Findings: Twenty-seven serologically defined malaria hotspots were detected in a survey conducted from 24 June to 31 July 2011 that included 17,503 individuals from 3,213 compounds in a 100-km2 area in Rachuonyo South District, Kenya. In a cluster-randomized trial from 22 March to 15 April 2012, we randomly allocated five clusters to hotspot-targeted interventions with larviciding, distribution of long-lasting insecticide-treated nets, indoor residual spraying, and focal mass drug administration (2,082 individuals in 432 compounds); five control clusters received malaria control following Kenyan national policy (2,468 individuals in 512 compounds). Our primary outcome measure was parasite prevalence in evaluation zones up to 500 m outside hotspots, determined by nested PCR (nPCR) at baseline and 8 wk (16 June–6 July 2012) and 16 wk (21 August–10 September 2012) post-intervention by technicians blinded to the intervention arm. Secondary outcome measures were parasite prevalence inside hotpots, parasite prevalence in the evaluation zone as a function of distance from the hotspot boundary, Anopheles mosquito density, mosquito breeding site productivity, malaria incidence by passive case detection, and the safety and acceptability of the interventions. Intervention coverage exceeded 87% for all interventions. Hotspot-targeted interventions did not result in a change in nPCR parasite prevalence outside hotspot boundaries (p ≥ 0.187). We observed an average reduction in nPCR parasite prevalence of 10.2% (95% CI −1.3 to 21.7%) inside hotspots 8 wk post-intervention that was statistically significant after adjustment for covariates (p = 0.024), but not 16 wk post-intervention (p = 0.265). We observed no statistically significant trend in the effect of the intervention on nPCR parasite prevalence in the evaluation zone in relation to distance from the hotspot boundary 8 wk (p = 0.27) or 16 wk post-intervention (p = 0.75). Thirty-six patients with clinical malaria confirmed by rapid diagnostic test could be located to intervention or control clusters, with no apparent difference between the study arms. In intervention clusters we caught an average of 1.14 female anophelines inside hotspots and 0.47 in evaluation zones; in control clusters we caught an average of 0.90 female anophelines inside hotspots and 0.50 in evaluation zones, with no apparent difference between study arms. Our trial was not powered to detect subtle effects of hotspot-targeted interventions nor designed to detect effects of interventions over multiple transmission seasons. Conclusions: Despite high coverage, the impact of interventions targeting malaria vectors and human infections on nPCR parasite prevalence was modest, transient, and restricted to the targeted hotspot areas. Our findings suggest that transmission may not primarily occur from hotspots to the surrounding areas and that areas with highly heterogeneous but widespread malaria transmission may currently benefit most from an untargeted community-wide approach. Hotspot-targeted approaches may have more validity in settings where human settlement is more nuclear. Trial registration: ClinicalTrials.gov NCT01575613.

opencc-zeroDec 2015View details →
zenodo32/100

Association of oligopaint FISH labeled super-enhancers and genes with polymerase clusters (raw data, SE1-SE6)

<p>This data set assesses the placement of different genomic regions relative to clusters formed by RNA polymerase II in zebrafish embryos. RNA polymerase was labeled by immunofluorescence, genomic regions by oligopaint DNA fluorescence in-situ hybridization (FISH). Microscopy images were acquired by instant-SIM microscopy, and analyzed using MatLab scripts and the bioformats importer. This data set contains the raw image data as well as all further analysis scripts.</p> <p>Raw image data for super enhancers SE1 to SE6</p>

opencc-by-4.0Jun 2021View details →
zenodo32/100

Microscopy of RNA Polymerase II clusters in live zebrafish embryos

<p>This data set contains raw image data and derived images illustrating clusters of RNA polymerase II as observed in zebrafish embryos. The clusters are labelled by fluorescently marked antibody fragments (Fab) that detect phosphorylation of the C-terminal roman of RNA polymerase II. Antibody fragments were provided by the laboratory of Hiroshi Kimura.</p>

opencc-by-4.0Jun 2021View details →
zenodo32/100

Association of oligopaint FISH labeled super-enhancers and genes with polymerase clusters (raw data, genes)

<p>This data set assesses the placement of different genomic regions relative to clusters formed by RNA polymerase II in zebrafish embryos. RNA polymerase was labeled by immunofluorescence, genomic regions by oligopaint DNA fluorescence in-situ hybridization (FISH). Microscopy images were acquired by instant-SIM microscopy, and analyzed using MatLab scripts and the bioformats importer. This data set contains the raw image data as well as all further analysis scripts.</p> <p>Raw image data for four genes: rnf19a, cdc25b, celf1, crsp7</p>

opencc-by-4.0Jun 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record