Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,326

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,326 results for “clusters”

Learn how ShareScore rates datasets ↗
dryad36/100

Data from: Genomic diversity of a nectar yeast clusters into metabolically, but not geographically, distinct lineages

Both dispersal limitation and environmental sorting can affect genetic variation in populations, but their contribution remains unclear, particularly in microbes. We sought to determine the contribution of geographic distance (as a proxy for dispersal limitation) and phenotypic traits (as a proxy for environmental sorting), including morphology, metabolic ability, and interspecific competitiveness, to the genotypic diversity in a nectar yeast species, Metschnikowia reukaufii. To measure genotypic diversity, we sequenced the genomes of 102 strains of M. reukaufii isolated from the floral nectar of hummingbird-pollinated shrub, Mimulus aurantiacus, along a 200-km coastline in California. Intraspecific genetic variation showed no detectable relationship with geographic distance, but could be grouped into three distinct lineages that correlated with metabolic ability and interspecific competitiveness. Despite ample evidence for strong competitive interactions within and among nectar yeasts, a full spectrum of the genotypic and phenotypic diversity observed across the 200-km coastline was represented even at a scale as small as 200 m. Furthermore, more competitive strains were not necessarily more abundant. These results suggest that dispersal limitation and environmental sorting might not fully explain intraspecific diversity in this microbe and highlight the need to also consider other ecological factors such as trade-offs, source-sink dynamics, and niche modification.

opencc-zeroDec 2017View details →
zenodo36/100

Globular Cluster Abundances from High-Resolution, Integrated-Light Spectroscopy. II. Expanding the Metallicity Range for Old Clusters and Updated Analysis Techniques

<p>Data from:</p> <p> Globular Cluster Abundances from High-Resolution, Integrated-Light Spectroscopy.<br>  II. Expanding the Metallicity Range for Old Clusters and Updated Analysis Techniques (Astrophysical Journal)</p> <p> J. E. Colucci, R. A. Bernstein, A. McWilliam, Observatories of the Carnegie Institution for Science</p> <p>This repository contains reduced globular cluster integrated light echelle spectra in IRAF readable format. <br> NOTE:  Spectra are *not* flux calibrated or doppler corrected. Sky/Background emission and absorption lines <br> are present. See reference paper for data reduction details.</p> <p>For each globular cluster:<br>  <br>  1.  *Approximately* normalized spectra are found in files ending with "ils_normalized.fits."  The echelle<br>  blaze function normalization was performed with an order by order fit to spectra of a reference G-type star.</p> <p> 2. Unnormalized spectra are found in files ending with "ils.fits." These spectra are not flux calibrated so do not<br>  use the count values in each order for science purposes. </p> <p><br> Spectra for the globular clusters NGC 104, NGC 362, NGC 2808, NGC 6093, NGC 6397, NGC 6752 were <br> taken with the DuPont telescope.  A reference star spectrum associated with the DuPont data is included : hr914_std.fits</p> <p>Spectra for the globular clusters NGC 6388, NGC 6440, NGC 6441, NGC 6528, NGC 6553 were taken with the <br> MIKE spectrograph on Magellan Clay.  A reference star spectrum associated with this data is included: ltt9239_std.fits</p> <p>Spectra for the globular cluster Fornax 3 was taken with the MIKE spectrograph on Magellan Clay on a different run. <br> A reference star spectrum associated with this data is included: hd033771_std.fits</p> <p>This research was supported by an NSF Astronomy and Astrophysics Postdoctoral Fellowship under award AST-1302710.</p>

opencc-by-4.0Oct 2016View details →
zenodo36/100

Tembetá from Abreu Garcia Cluster 8 Cremated Deposit RTI First Release

<p>First release of RTI file for stone tembetá (lip-plug) recovered from Cluster 8 cremated deposit of Mound A, at the Abreu Garcia Mound and Enclosure complex, Santa Catarina, Brazil during the 2015 excavation season as part of the Jê Landscapes of Southern Brazil project. Created with 13 images using RTIBuilder 2.0.2 and the Hemispherical Harmonics (HSH) fitter.</p> <p>Primeira divulgação do arquivo de RTI de um adorno labial em pedra - tembetá recuperado do depósito de remanescentes cremados - Cluster 8 no Mound A, no MEC Abreu &amp; Garcia, Santa Catarina, Brasil, durante a campanha de escavação de 2015, como parte do projeto Paisagens Jê do Sul do Brasil. Criado com 13 imagens usando RTIBuilder 2.0.2 e Hemispherical Harmonics (HSH) fitter.</p>

opencc-by-sa-4.0Apr 2017View details →
zenodo36/100

Data for Properties of Kinetic Transition Networks for Atomic Clusters and Glassy Solids

<p>Databases of minima and transition states for Morse clusters, in two and three dimensions at a variety of ranges.</p>

opencc-by-4.0May 2017View details →
zenodo36/100

USEARCH Clustering Data for the Influenza Internal Genes

<p>These are the aligned sequence clusters and centroid datasets for the USEARCH clustering of the influenza A internal genes. </p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

Study protocol and data dictionary: Effectiveness of a GP delivered medication review in reducing polypharmacy and potentially inappropriate prescribing in older patients with multimorbidity in Irish primary care: a cluster randomised controlled trial (SPPiRE study)

<p><strong>Methods</strong></p> <p><strong>Study design and participants</strong></p> <p>The methods for the SPPiRE cluster RCT have been described in the trial protocol (21). This study is reported in line with the CONSORT 2010 cluster RCT checklist (22), see Appendix 1, and was approved by the Irish College of General Practitioners Research Ethics Committee. In brief, SPPiRE was a pragmatic two arm cluster RCT, with the intervention delivered to GP clusters and analysis of outcomes at the patient level. Information about the trial was publicised through a variety of GP research, teaching and training networks throughout Ireland. Eligible practices expressing an interest were formally invited. Practices were eligible to participate if they had at least 300 registered patients aged &ge;65 years (based on the need to identify a sufficient number of eligible participants) and used either of the two Irish GP practice management systems (PMS) with over 80% national cover; this enabled use of a SPPiRE patient finder tool which was developed and embedded into these systems. Practices were excluded if they were currently involved in a medication management or prescribing trial or if they were unable to recruit at least five participants.</p> <p>Eligible patients were aged &ge;65 years and prescribed &ge;15 repeat medicines. A repeat medicine was defined as any unique item with a World Health Organisation Anatomical Therapeutic Chemical code on the patient&rsquo;s current repeat prescription. Patients were excluded if they had been recruited into a practice that was unable to recruit at least four other participants, they were judged by their GP as unable to give informed consent or they were unable to attend the practice for a face to face medication review, (e.g. nursing home residents and house bound patients).&nbsp; Recruited GPs ran the SPPiRE patient finder tool and screened the generated list to ensure only eligible patients were invited. Practices who identified more than 40 eligible patients were supported in selecting a random sample of 30 patients to invite. All recruited practices and patients gave fully informed consent and baseline data was collected prior to practice allocation, to reduce the likelihood of selection bias.</p> <p><strong>Randomisation and masking</strong></p> <p>Recruited practices were allocated to intervention or control groups by minimisation using Minimpy software (23) by the trial statistician (FB) who had no knowledge of participating practices. Minimisation variables included practice size (number of GP sessions per week, 0-14, 14-28 and 28 or more) and location (urban, rural or mixed). Considering the nature of the intervention, it was not possible to blind GPs or patients to the intervention, however to reduce the risk of detection bias the two primary outcome measures; the number of repeat medicines and whether a PIP was present were assessed by an independent blinded pharmacist (MF).</p> <p><strong>Procedures</strong></p> <p>Intervention GPs received unique login details to the SPPiRE website where they had access to five training videos and a template for performing the SPPiRE medication review. The training videos provided background information on multimorbidity and polypharmacy, PIP, eliciting patient treatment priorities and conducting a brown bag medication review. GPs were instructed to book a double appointment and to ask their patients to bring all their medicines in to the medication review visit with them. The SPPiRE medication review process had two main components; gather and record information and then to discuss and agree changes with their patient based on the recorded information, with a focus on deprescribing medicines that were potentially inappropriate, figure 1. The website provided suggested treatment alternatives for identified PIP but all treatment decisions were ultimately at the discretion of the individual GP, based on their clinical judgement and their patients&rsquo; individual priorities.</p> <p>Control GPs delivered usual care during the six to twelve month study period. At the time of intervention delivery there was no structured chronic disease management programme in Irish primary care and many patients with multimorbidity attended multiple hospital specialists. In Ireland, the majority of people aged &ge;70 years of age have access to free GP visits and medicines with some prescription charge co-payments. In the 65 &ndash; 69 year old age category a lower proportion have access to both free GP visits and prescription medicines. Access to specialists and diagnostics in secondary care is free for the entire population.</p> <p><strong>Outcomes</strong></p> <p>The two primary outcomes were the number of repeat medicines and the proportion of patients with any PIP, from a list of 34 pre-specified indicators (see Appendix 2). A series of secondary prescribing related outcome measure were pre-specified to allow a more in depth analysis of the effect of the intervention on prescribing. These were:</p> <ul> <li>The number of medicines stopped and started</li> <li>The proportion of patients with a reduction in significant polypharmacy (defined as &ge;15 repeat medicines)</li> <li>The number of PIP</li> <li>The proportion of patients with a high risk PIP (see Appendix 2)</li> <li>The proportion of patients with any reduction in PIP</li> </ul> <p>Secondary patient reported outcomes measures were included to capture the effectiveness of the intervention from the patients&rsquo; perspective. These were:</p> <ul> <li>Health related Quality of life (EQ5D-5L)(24)</li> <li>Revised Patients' attitudes towards deprescribing (rPATD)&nbsp;&nbsp;(25)</li> <li>Multimorbidity Treatment Burden Questionnaire (MTBQ)&nbsp;(26)</li> </ul> <p>Health care utilisation data was collected to assess the effect of the intervention on health care usage and for the trial&rsquo;s economic evaluation.</p> <p>Outcomes were collected at baseline and at six months after intervention delivery. Patient reported measures were collected by postal questionnaires. Data for all other measures including prescribed medicines, medical and investigations history and healthcare utilisation were collected by participating GPs and submitted to the study manager (CMC). This was a deviation from the original protocol, which indicated this data would be collected by the research team. This deviation related to changes in data protection and national health research regulations during the study period, which precluded research team access to the patients&rsquo; full clinical record.</p> <p><strong>Adverse events</strong></p> <p>Information on adverse events such as mortality, ED presentations and hospital admissions was collected at follow up. Given the deprescribing approach of the intervention a safety protocol for identifying and reporting any suspected adverse drug withdrawal events (ADWEs) was developed. An ADWE is defined as either recurrence of the condition for which the drug was prescribed (e.g. recurrence of angina after stopping a beta blocker) or a physiologic reaction to drug withdrawal (e.g. SSRI withdrawal syndrome)&nbsp;&nbsp;(27, 28). Although discontinuing medicines in older people has been demonstrated to be safe (29), given the paramount importance of the principle of &ldquo;do no harm&rdquo; in research ethics a vigorous and detailed method was established to ensure that any potential ADWEs precipitated by deprescribing in a SPPiRE medication review were captured. Intervention GPs were asked to report any possible ADWE following the SPPiRE medication review. The Naranjo ADR probability scale (30) has been adapted in other studies to assess the likelihood a reaction is related to drug withdrawal&nbsp;&nbsp;(27, 28). This tool was further adapted for SPPiRE and used to make an assessment on the causality of the ADWE. To ensure the patient perspective was included, self-reported possible ADWEs were also collected from patient follow up questionnaires.&nbsp;&nbsp;</p> <p>&nbsp;</p> <p><strong>Sample size</strong></p> <p>As outlined in the trial protocol (21), the study was designed with 90% power to detect a 20% reduction in the proportion with PIP and a mean difference of one medicine between intervention and control groups (based on a mean of 17.4 medicines SD (2.6)) and the sample size inflated to incorporate the effects of clustering (using an ICC of 0.025). The sample size was recalculated when it became apparent during early recruitment that it would not be possible to recruit clusters with an average of 15 participants, as was initially planned in the protocol. An average cluster size of eight was anticipated which inflated the original sample size from 30 practices (450 patients) to 50 practices (400 patients).</p> <p><strong>Statistical analysis</strong></p> <p>Descriptive statistics were used to describe baseline characteristics of recruited practices and participants. All analyses were conducted under the intention-to-treat principle and those lost to follow up had their baseline data carried forward. The primary analysis was carried out using multi-level modelling. The first primary outcome measure, number of repeat medications, was assessed using mixed effects Poisson regression with the individual as the unit of analysis and the practice included as the random effect to control for the effects of clustering and results presented using incidence rate ratios (IRR) and 95% confidence intervals (CI). The baseline number of medicines, GP size (number of GP sessions per week) and GP location (urban/rural) were included in the analysis as fixed effects. The second outcome measure, proportion of patients with a PIP, was analysed in a similar manner using mixed effects logistic regression, including PIP at baseline, GP size and location, and results presented using odd ratios (OR) and 95% CIs. A number of pre-specified sensitivity analyses were conducted; complete case analysis, per protocol analysis and including &ldquo;presence of a repeat prescribing policy&rdquo; as a covariate. All secondary outcomes were analysed in a similar manner to the primary outcomes, using appropriate mixed effects regression methods (i.e. linear, logistic, Poisson).</p> <p>&nbsp;</p> <p>Note: Version 3 (published 28 April 2025) updates Version 2 by removing Participant GP1P4 following consent withdrawal. This version should be used for all future analyses.</p> <p>&nbsp;</p>

opencc-by-4.0May 2021View details →
zenodo36/100

Supplemental data for "The genome of a sea spider corroborates a shared Hox cluster motif in arthropods with reduced posterior tagma"

<div># List of files on Zenodo</div> <p>&nbsp;</p> <div>Data accompanying the manuscript have been uploaded on Zenodo (10.5281/zenodo.14185694). Here we</div> <div>present a brief description of each file and put them in meaningful groups.</div> <p>&nbsp;</p> <div>## referenced supplement</div> <p>&nbsp;</p> <div>(Cited) supplementary material from the manuscript. For more information, please refer to the figure/table legends and the supplementary file descriptions.</div> <p>&nbsp;</p> <div> <div>- add-file-01.pdf</div> <div>- add-file-02-table1-data_overview.tsv</div> <div>- add-file-03-table2-genome_progress.tsv</div> <div>- add-file-04-table3-Pycnognonum_microRNAs.tsv</div> <div>- add-file-05-table4-named_genes.tsv</div> <div>- add-file-06-table5-abdA.tsv</div> <div>- add-file-07-hox_tree.pdf</div> <div>- add-file-08-table6-r2_g3735-Alignment-HitTable.tsv</div> <div>- add-file-09-hro_tree.pdf</div> <div>- add-file-10-irx_tree.pdf</div> <div>- add-file-11-sine_tree.pdf</div> <div>- add-file-12-nk_tree.pdf</div> <div>- add-file-13-dbx.png</div> <div>- add-file-14-alignment.pdf</div> <div>- add-file-15-Plit_COI-alignment_distances.pdf</div> <div>- add-file-16-table7-isoseq.tsv</div> <div>- add-file-17-table8-chelicerate_repeat_content.tsv</div> <div>- add-file-18-table9-arthropod_genomes.tsv</div> <div>- add-file-19-table10-arthropod_repeat_content.tsv</div> <div>- add-file-20-table11-chelicerate_genomes.tsv</div> <div>- add-file-21-gene_analysis.zip</div> <div>- add-file-22-table12-self_synteny.tsv</div> </div> <p>&nbsp;</p> <div>## figures</div> <p>&nbsp;</p> <div>- figs.zip: archive of raw and processed figures for the manuscript in full resolution</div> <p>&nbsp;</p> <div>## analysis</div> <p>&nbsp;</p> <div>### genomic context</div> <p>&nbsp;</p> <div>The broader arthropod/chelicerate context for the _P. litorale_ genome assembly.</div> <p>&nbsp;</p> <div>- araneae.tsv: repeat makeup of published chelicerate assemblies</div> <div>- arthropoda.tsv: genome assembly statistics for arthropod genomes. From NCBI Genomes.</div> <div>- modern_taxids.txt: list of taxonomic IDs for species; made to be submitted to NCBI Taxonomy.</div> <div>- tax_report.txt: the full taxonomic report for each query species. Contains tax IDs for the entire lineage.</div> <div>- total_repeats.tsv: total repeat content of published chelicerate genomes.</div> <p>&nbsp;</p> <div>### Homeobox genes</div> <p>&nbsp;</p> <div>Files concerning the Homeobox gene cluster analysis. Each folder (hro, irx, nkx, sine) contains the</div> <div>candidate sequences from Aase-Remedios et al., the P. litorale sequences that matched, the multiple</div> <div>sequence alignment, the trimmed alignment, and the tree files.</div> <p>&nbsp;</p> <div>Additionally, the hox/ folder contains the analysis done for the AbdA gene, with searches performed</div> <div>against the de-novo assembled transcriptomes (.m8 files).</div> <p>&nbsp;</p> <div>Finally, the r2_g3735/ folder contains the analysis of the r2_3735 gene model, which is found in the</div> <div>Hox cluster area on the P. litorale genome. Sequence searches against NCBI nr and the developmental</div> <div>transcriptomes show that the gene has putative homologs in other taxa and is expressed throughout</div> <div>development.</div> <p>&nbsp;</p> <div>## processed (intermediate) data</div> <p>&nbsp;</p> <div>### 00-kmer-jellyfish.zip</div> <p>&nbsp;</p> <div>k-mer spectra analysis with GenomeScope and GenomeScope2.0</div> <p>&nbsp;</p> <div>### 00-seq-qc.zip</div> <p>&nbsp;</p> <div>quality control output for raw sequencing data (e.g. FastQC output)</div> <p>&nbsp;</p> <div>### 01-assembly</div> <p>&nbsp;</p> <div>- assembly_graph.gfa: Flye output</div> <div>- assembly_graph.gv: Flye output</div> <div>- assembly_info.txt: Flye output</div> <div>- assembly.fasta: Flye output</div> <div>- backmap.hifi.sort.bam.cov-hist.pdf: coverage histogram of the back-mapped PacBio data</div> <div>- backmap.ont.sort.bam.cov-hist.pdf: coverage histogram of the back-mapped ONT data</div> <div>- BUSCO.arthropoda_odb10.txt: BUSCO completeness report (arthropoda_odb10)</div> <div>- BUSCO.metazoa_odb10.txt: BUSCO completeness report (metazoa_odb10)</div> <div>- flye.log: Flye assembler log</div> <div>- quast_report.pdf: assembly QC</div> <p>&nbsp;</p> <div>### 02-scaffold</div> <p>&nbsp;</p> <div>`yahs` output files:</div> <p>&nbsp;</p> <div>- asm_hic.sorted.bam</div> <div>- flye-yahs.fa</div> <div>- yahs.out_scaffolds_final.agp</div> <div>- yahs.out_scaffolds_final.fa.hic</div> <div>- yahs.out_scaffolds_final.fa.assembly</div> <p>&nbsp;</p> <div>Juicebox (manual curating) results:</div> <p>&nbsp;</p> <div>- yahs.out_scaffolds_final.fa.review.assembly</div> <div>- 02-flye-yahs-juicebox.fa</div> <p>&nbsp;</p> <div>GAP `sort_scaffolds` pipeline outputs</div> <p>&nbsp;</p> <div>- 03-flye-yahs-juicebox-merge.fasta</div> <div>- plit_q_0_50000_0.5FracBest_unseen_scaffolds.txt</div> <div>- plit_q_0_50000_0.5FracBest_insertion_stats.tsv</div> <div>- plit_q_0_50000_0.5FracBest_appended_scaffolds.tsv</div> <div>- plit_q_0_50000_0.5FracBest_inserted_scaffolds.tsv</div> <p>&nbsp;</p> <div>### 03-contamination</div> <p>&nbsp;</p> <div>Refer to the [contamination analysis](https://github.com/galicae/plit-genome/blob/main/04-contam/README.md) for details.</div> <p>&nbsp;</p> <div>Decontaminating the draft genome from non-metazoan scaffolds:</div> <p>&nbsp;</p> <div>- plit_q_0_50000_0.5FracBest_output_filtered.fasta: input draft genome</div> <div>- contam_tax.m8: Alignment results of UniRef90 against draft genome (MMseqs2)</div> <div>- scaffolds_taxonomic_distribution.tsv: summary of contam_tax.m8; number of genes from each taxonomic level per scaffold.</div> <div>- scaffolds_taxonomic_distribution_collapsed_vir.tsv: scaffolds with predominantly viral hits</div> <div>- scaffolds_taxonomic_distribution_suspect.tsv: scaffolds whose genes are &lt;90% metazoan</div> <p>&nbsp;</p> <div>Checking for widespread _Metridium_ contamination:</div> <p>&nbsp;</p> <div>- primary_mq30.txt: list of high-quality mapping reads (presumptive "metridial")</div> <div>- metridium_scaffolds.txt_summary: no. of presumptive _Metridium_ reads per draft scaffold</div> <div>- metridium_scaffolds.txt: filtered SAM file with all high-quality "_Metridium_" hits on draft scaffolds</div> <div>- metridium_contigs.sam_summary: no. of presumptive _Metridium_ reads per Flye contig</div> <p>&nbsp;</p> <div>### 04-annotation</div> <p>&nbsp;</p> <div>Repeat analysis with RepeatModeler/RepeatMasker:</div> <p>&nbsp;</p> <div>- draft.fasta.tbl: output of RepeatModeler in tabular form</div> <div>- pb.sam.flagstats: summary of mapping the repeat families to the PacBio data.</div> <div>- pycno-families.fa: output of RepeatModeler - the sequences of the _P. litorale_ repeat families</div> <div>- draft.fasta.out.gff: output of RepeatModeler - repeat locations on the draft genome</div> <p>&nbsp;</p> <div>Protein coding gene annotation:</div> <p>&nbsp;</p> <div>- annot-01-isoseq.gff: GFF file with the gene models proposed using Iso-seq isoforms</div> <div>- annot-01-braker.gff: GFF file with the gene models proposed from round 1 of BRAKER3 using developmental transcriptomes</div> <div>- annot-02-braker.gff3: GFF file with the gene models proposed from round 2 of BRAKER3, using developmental transcriptome reads that weren't used in round 1</div> <div>- annot-03-denovo.gff3: GFF file with the gene models proposed from the de novo transcriptomes</div> <div>- deep_denovo_assemblies.zip: the de-novo assembled transcriptomes from the deeply sequenced developmental time points. Also available on ENA.</div> <p>&nbsp;</p> <div>tRNAscan output</div> <p>&nbsp;</p> <div>- trnascan.bed</div> <div>- trnascan.out</div> <div>- trnascan.fasta</div> <div>- trnascan.stats</div> <p>&nbsp;</p> <div>MirMachine output</div> <p>&nbsp;</p> <div>- Pli_september.PRE.gff: MirMachine output with permissive threshold</div> <div>- Pli_september.PRE-1.gff: MirMachine output with strict threshold</div> <div>- Pli_september.PRE.fasta: predicted miRNA sequences</div> <p>&nbsp;</p> <div>## Results</div> <p>&nbsp;</p> <div>- draft_softmasked.fasta: draft genome with repetitive regions softmasked</div> <div>- draft.fasta: draft genome fasta</div> <div>- hox.gff3: the position of the Hox genes in GFF3 form.</div> <div>- merged_sorted_named_dedup_flagged.gff3: protein-coding gene models from all rounds of annotation after deduplication</div> <div>- transcripts.fa: TransDecoder-extracted putative transcripts</div> <div>- transcripts.fa.transdecoder.pep: TransDecoder predicted peptides</div> <div>- out.emapper.annotations: EggNOG-mapper functional annotation for the predicted peptides</div> <div>- out.emapper.best.annotations: filtered EggNOG-mapper annotation; best hit per gene kept</div> <div>- miRNA.fasta: FASTA sequences of predicted miRNAs</div> <div>- miRNA.lenient.gff: GFF of miRNA positions (permissive MirMachine cutoff)</div> <div>- miRNA.strict.gff: GFF of miRNA positions (strict MirMachine cutoff)</div>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Data supporting "Local Clustering Decoder: a fast and adaptive hardware decoder for the surface code"

<p>The data consists of a CSV file containing the raw performance data collected from running our decoder on a Xilinx Virtex Ultrascale+ VU19P FPGA. The Stim circuits that were used to create the samples are provided in a ZIP file. Our internal fork of Stim with support for leakage is needed to sample noise from the circuits.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Effects of Type Ia Supernovae in young globular clusters

<p>Talk at Elba 23, conference in honor of Mike Rich</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Dalk_Glacier_clustering_2023

<p>Dalk_clustering_events.tar.gz contains 909,720 events' waveforms detected&nbsp;near Dalk Glacier, in the Larsemann Hills, East Antarctica&nbsp;from 6 Dec 2019 to 2 Jan 2020. Original seismic data are recorded by 100 three-component short-duration seismometers. The preprocessing includes detrending, tapering and 1-Hz high-pass filtering and the detection method is&nbsp;STA/LTA (short term=0.5 s, long term=30 s, trigger threshold=10, detrigger threshold=3).</p><p><br>AWS_environment.zip contains recordings of wind speed, wind direction and temperature at the Zhongshan Station from 6 Dec 2019 to 2 Jan 2020.</p><p>unsupervised_clustering_code.zip contains the python codes of detecting by STA/LTA, feature extraction by an autoencoder and clustering by a Gaussian mixture model.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

POMFinder: Identifying polyoxometalate cluster structures from pair distribution function data using explainable machine learning

<p>Databases that were used to train the POMFinder ML model.</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Synthesis of urea on the surface of interstellar water ice clusters.

<p>The&nbsp;Supporting Material contains the cartesian coordinates of B3LYP-D3(BJ)/ma-def2-TZVP&nbsp;optimized minima and transition state for the reaction studied in the paper, in .xyz format, computed using&nbsp;<a href="https://gaussian.com/">O</a><a href="https://orcaforum.kofo.mpg.de/">RCA</a>&nbsp;code. The folders are relative to:</p> <ul> <li>gas phase reactions;</li> <li>reactions modelled on the 18 water molecules cluster: <ul> <li>closed-shell +&nbsp;closed-shell;</li> <li>closed-shell + radical;</li> <li>radical + radical (2 outcomes);</li> <li>ion-pair formation (2 pathways);</li> </ul> </li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Athens Biodiversity Clustering Dataset - Features and Clusters per region

<p>Athens Biodiversity Clustering Dataset - Features and Clusters per region. Includes collated data from a variety of sources, including local data and remote sensing data. All data is high resolution, with features available for 490 census blocks in Central Athens.</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Atomic Cluster Expansion for a General-Purpose Interatomic Potential of Magnesium

<p>This collection contains files associated with Physical Review Materials. "Atomic cluster expansion for a general-purpose interatomic potential of magnesium" (2023) paper:</p><p>- ACE potentials for magnesium.</p><p>-Active set inverted (ASI) for the ACE potential</p><p>- Magnesium DFT-PBE dataset computed with FHI-aims and that was used for fitting Atomic Cluster Expansion potential for magnesium.</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Reproduction Package for the paper "The early evolution of young massive clusters. II. The kinematic history of NGC 6618 / M 17"

<h2>Reproduction package for the paper "The early evolution of young massive clusters. II. The kinematic history of NGC 6618 / M 17".</h2><ul><li>This reproduction package aims for open science, with the internal API designation of 'Gold'</li><li>Authors: M. Stoop, A. Derkink, L. Kaper, A. de Koter, C. Rogers, M.C. Ramírez-Tannus, D. Guo, N. Azatyan</li><li>Paper DOI: https://doi.org/10.1051/0004-6361/202347383</li><li>Arxiv DOI: https://doi.org/10.48550/arXiv.2311.04174</li><li>Zenodo DOI: http://doi.org/10.5281/zenodo.8120575</li><li>Accepted for publication in Astronomy and Astrophysics (date of acceptance: 2023/10/27)</li></ul><h2>Raw Data</h2><ul><li>./gaia_files/ gives the ADQL queries on how to obtain the raw data from the Gaia database along with the raw Gaia data itself.</li><li>./other_files/ gives the other raw data obtained from other sources. For example the literature, Aladin, parsec isochrones.</li><li>We have cited the relevant references in the paper for raw data taken from the literature.</li><li>Optical VLT, SALT, WHT spectra (raw data, software and end-products) described in Appendix A can be requested from Annelotte Derkink.</li></ul><h2>Software</h2><ul><li>MacOS Big Sur 11.6</li><li>Jupyter Notebook (6.3.0)</li><li>Programming languages used: Python (3.9.7)</li><li>Python packages used: numpy (1.22.3), math (comes with Python), pandas (1.2.1), matplotlib (3.3.3), scipy (1.6.0), os (comes with Python), zero_point (0.0.1) (https://gitlab.com/icc-ub/public/gaiadr3_zeropoint), pyUPMASK (requires os, time, multiprocessing, astropy, pathlib) (https://github.com/msolpera/pyUPMASK)</li><li>Aladin Desktop (Version 11.0) (https://aladin.u-strasbg.fr/AladinDesktop/), will need java installed</li><li>Aladin Lite (https://aladin.cds.unistra.fr/AladinLite/)</li></ul><h2>Figures and Tables</h2><ul><li>Figures can be reproduced from the ./figures/ folder.</li><li>All material and data used are available either in the Raw Data or in the Intermediate Data</li><li>Jupyter notebooks (.ipynb files) provide the option to execute all code as ready-made to produce the figures: 'Cell' -&gt; 'Run All'.</li><li>The figures shown in the paper will be saved in ./figures/figures_paper/ folder.</li><li>Tables can be reproduced from the ./tables/ folder similar to the figures.</li><li>Tables are saved in .csv files so they can easily be used. Tables are also saved in _latex.csv form for easy implementation in latex.</li></ul><h2>Intermediate data products</h2><ul><li>Intermediate data, such as the determined members of NGC6618 or found runaway stars, can be found in the /output_files/ folder.</li><li>The 'raw' data is also given here, because small corrections have been applied, which are recommended by the Gaia Consortium. This 'raw' data file is not the same as the one given in the ./gaia_files/ folder! This is to ensure that both the raw and intermediate data products are available.</li><li>./modules/ is part of the Python pyUPMASK module.</li></ul><h2>End-to-End analysis scripts</h2><ul><li>The end-to-end analysis to determine the members is given in the Jupyter Notebook 'membership_pyupmask.ipynb'.</li><li>The end-to-end analysis to search for runaways is given in the Jupyter Notebook 'runaways.ipynb'.</li><li>Minor 'end-to-end' analysis is done in the relevant figure or table Jupyter Notebook.</li></ul>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Figures for: Biosynthetic gene clusters carried by plasmids may enhance adaptations to a changing ocean

<p>Figures created for the short communication: Biosynthetic gene clusters carried by plasmids may enhance adaptations to a changing ocean.</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Pairing Remote Sensing and Clustering in Landscape Hydrology for Large-Scale Changes Identification. Applications to the Subarctic Watershed of the George River (Nunavik, Canada). Dataset and Code.

<p>For remote and vast northern watersheds, hydrological data are often sparse and incomplete. Landscape hydrology provides useful approaches for the indirect assessment of the hydrological characteristics of watersheds through analysis of landscape properties. In this study, we used unsupervised Geographic Object-Based Image Analysis (GeOBIA) paired with the Fuzzy C-Means (FCM) clustering algorithm to produce seven high-resolution territorial classifications of key remotely sensed hydro-geomorphic metrics for the 1985-2019 time-period, each spanning five years. Our study site is the George River watershed (GRW), a 42,000 km<sup>2</sup> watershed located in Nunavik, northern Quebec (Canada). The subwatersheds within the GRW, used as the objects of the GeOBIA, were classified as a function of their hydrological similarities. Classification results for the period 2015-2019 showed that the GRW is composed of two main types of subwatersheds distributed along a latitudinal gradient, which indicates broad-scale differences in hydrological regimes and water balances across the GRW. Six classifications were computed for the period 1985-2014 to investigate past changes in hydrological regime. The seven-classification time series showed a homogenization of subwatershed types associated to increases in vegetation productivity and in water content<br> in soil and vegetation, mostly concentrated in the northern half of the GRW, which were the major changes occurring in the land cover metrics of the GRW. An increase in vegetation productivity likely contributed to an augmentation in evapotranspiration and may be a primary driver of fundamental shifts in the GRW water balance, potentially explaining a measured decline of about 1 % (&sim; 0.16 km<sup>3</sup>y<sup>&minus;1</sup>) in the George River&rsquo;s discharge since the mid-1970s. Permafrost degradation over the study period also likely affected the hydrological regime and water balance of the GRW. However, the shifts in permafrost extent and active layer thickness remain difficult to detect using remote sensing based approaches, particularly in areas of discontinuous and sporadic permafrost.</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Differentially expressed genes in berries and rachis of berry shrivel grape clusters used to prepare figures for a review

<p>Grapevine berry shrivel is a ripening disorder leading to significant economic losses in the worldwide wine and table grape industries. Sugar accumulation stops early after ripening onset accompanied with cell death in berreis and subtending pecicels and rachis finally resulting in berry shrinkage. To date, the triggers of BS remain unknown. The dataset supports figures prepared for an review which aims to summarize and critically discuss the current knowledge. Data are expressed as differentially expressed genes obtained from grape berries samples collected at six developmental stages (pre- until post-veraison) analysed with RNASeq and two pooled samples (pre- and BS symptomatic) from the rachis analyzed with a microarray study. Extracted information focus on primary metabolic processes including sugar transport and metabolism, organic acid metabolism, stress signaling and cell as well as cell wall organisation. Data are mean values of three biological represent and presented as log2 fold changes including statistical information. A meta-data sheet provides the most relevant information and references.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Effect of Digoxin on clusters of circulating tumor cells in patients with metastatic breast cancer: a phase 1 trial

<p>This repository contains processed transcriptomics data, large data sets and additional files required to reproduce the code available at the repository https://github.com/TheAcetoLab/dicct-trial</p>

opencc-by-4.0Dec 2025View details →
zenodo36/100

Data presented in figures of "Measurement Report: Insights into the chemical composition and origin of molecular clusters and potential precursor molecules present in the free troposphere over the Southern Indian Ocean: observations from the Maïdo observatory (2150 m a.s.l., Reunion Island)"

<p>This dataset includes the data shown in the figures of "Measurement Report: Insights into the chemical composition and origin of molecular clusters present in the free troposphere over the Southern Indian Ocean: observations from the Maïdo observatory (2150 m a.s.l., Reunion Island)". Read me files containing information on the reported data can be found in the different folders.&nbsp;</p>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record