Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,481
datasets available to search
ShareScore release 0.9.0
Dataset results
1,481 results for “data processing”
Cascade Project at North Temperate Lakes LTER Core Data Process Data 1984 - 2016
Data useful for calculating and evaluating primary production processes were collected from 6 lakes from 1984-2016. Chlorophyll a and pheophytin were measured by the same fluorometric method from 1984-2016. In some years chlorophyll and pheophytin were separated into size fractions (total, and a 'small' fraction that passed a 35 um mesh screen). Primary production was measured by the 14C method from 1984-1998. Dissolved inorganic carbon for primary production calculation was calculated from Gran alkalinity titration and air-equilibrated pH until 1987 when this method was replaced by gas chromatography. Until 1995 alkaline phosphatase activity was measured as an indicator of phosphorus deficiency.
MCR LTER: Coral Reef: Dead coral skeletons impair key recovery processes following coral bleaching; data for Kopecky et al., 2024 Global Change Biology
The data included in this data package were collected on the North shore of Moorea, French Polynesia, from 2015-2023 to explore how dead coral skeletons (e.g,, left after coral bleaching events) influence critical processes tied to coral reef resilience. Together, these various datasets were used for analyses in the manuscript entitled "Changing disturbance regimes, material legacies, and stabilizing feedbacks: dead coral skeletons impair key recovery processes following coral bleaching", published in Global Change Biology. These data are in support of a publication Kopecky et al. (2024) Global Change Biology, and were a part of the thesis of K. Kopecky. The manuscript title and author list are as follows: Changing disturbance regimes, material legacies, and stabilizing feedbacks: dead coral skeletons impair key recovery processes following coral bleaching. Kai Kopecky, Russell J. Schmitt, Sally J. Holbrook. This material is based upon work supported by the U.S. National Science Foundation under Grant No. OCE 22-24354 (and earlier awards) as well as a generous gift from the Gordon and Betty Moore Foundation. Research was completed under permits issued by the French Polynesian Government (Délégation à la Recherche) and the Haut-commissariat de la République en Polynésie Francaise (DTRT) (Protocole d'Accueil 2005-2024). This work represents a contribution of the Moorea Coral Reef (MCR) LTER Site.
Cascade Project at North Temperate Lakes LTER: Process Data 1984 - 2007
Data on chlorophyll, primary productivity, and alkaline phosphatase activity from 1984-95. Samples were collected with a Van Dorn bottle at 6 depths determined from the percent of surface irradiance (100%, 50%, 25%, 10%, 5% and 1%) and in the hypolimnion (12 m in Peter, East Long, West Long, and Tuesday lakes; 9 m in Paul Lake; and 4.5 m in Central Long Lake). Sampling Frequency: varies Number of sites: 8
Dominant contribution of Asgard archaea to eukaryogenesis (2024) Tobiasson, V., Koonin, E. PROCESSED DATA AND METADATA
<h1>Main data deposit for "Dominant contribution of Asgard archaea to eukaryogenesis". </h1> <p>Victor Tobiasson, Jacob Luo, Yuri I Wolf, Eugene V Koonin</p> <p>Computational Biology Branch, Division of Intramural Research, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA</p> <p><strong>The Origin of eukaryotes is one of the key problems in evolutionary biology. The demonstration that the Last Eukaryotic Common Ancestor (LECA) already contained the mitochondrion, an endosymbiotic organelle derived from an alphaproteobacterium, and the discovery of Asgard archaea, the closest archaeal relatives of eukaryotes inform and constrain evolutionary scenarios of eukaryogenesis. We undertook a comprehensive analysis of the origins of the core eukaryotic genes tracing to the LECA within a rigorous statistical framework centered around evolutionary hypotheses testing using constrained phylogenetic trees. The results reveal dominant contributions of Asgard archaea to the origin of most of the conserved eukaryotic functional systems and pathways. A limited contribution from Alphaproteobacteria was identified, primarily relating to the energy transformation systems and Fe-S cluster biogenesis, whereas ancestry from other bacterial phyla was scattered across the eukaryotic functional landscape, without consistent trends. These findings suggest a model of eukaryogenesis in which key features of eukaryotic cell organization evolved in the Asgard ancestor, followed by the capture of the Alphaproteobacterial endosymbiont, and augmented by numerous but sporadic horizontal acquisition of genes from other bacteria both before and after endosymbiosis. </strong></p> <div> <div> <div>Version 0.3, updated 180325</div> <div> </div> <div> </div> <div>Main data repository for:</div> <div>Dominant contribution of Asgard archaea to eukaryogenesis (2024) </div> <div>Tobiasson, V., Koonin, E.</div> <div> </div> <div>Contains all final parsed data from the main Eukaryogenesis project </div> <div>investigating the evolutionary ancetries of eukaryotic protein families. </div> <div> </div> <div>Currently (non-static) available at: </div> <div>https://www.biorxiv.org/content/10.1101/2024.10.14.618318v2</div> <div>https://assets-eu.researchsquare.com/files/rs-5352492/v1/2f9c68ae-cf3e-420a-8d29-867b6fb1a878.pdf</div> <div> </div> <div>All code used to generate the data present within this repository available at: </div> <div>https://github.com/VictorTobiasson/eukgen </div> <div> </div> <div> </div> <div>### General information</div> <div> </div> <div>To identify associations between prokaryotic and eukaryotic protein families, separate</div> <div>hidden Markov model (HMM) databases for prokaryotes and eukaryotes were constructed </div> <div>using a custom, cascaded, sequence-to-profile clustering pipeline, implemented using </div> <div>mmseqs2, followed by a multistep data-reduction and multiple sequence alignment (MSA) </div> <div>procedure to generate HMM profiles using hhsuite. </div> <div> </div> <div>A prokaryotic database of 37 million protein sequences was curated from prokaryotic </div> <div>genomes obtained from the NCBI GenBank in November 2023 and supplemented with proteins </div> <div>extracted from 146 Asgard genome assemblies. To avoid inclusion of genes present only </div> <div>within a narrow subset of species, possibly resulting from horizontal transfer from </div> <div>eukaryotes post LECA, we reconstructed the “soft-core” pangenome for each of the 26 </div> <div>curated prokaryotic taxonomic classes. These pangenomes include only those genes that </div> <div>are present in at least 67% of the families within each class of Bacteria and Archaea. </div> <div>The initial eukaryotic database consisted of 30 million protein sequences from 993 </div> <div>species taken from EukprotV3 and cleaned using mmseqs2 to remove likely prokaryotic </div> <div>contaminants. </div> <div> </div> <div>Both databases were clustered and MSAs constructed for all non, singleton clusters </div> <div>and HMM profiles created. The resulting eukaryotic HMM dataset was queried against </div> <div>the prokaryotic dataset using hhblits to identify sets of homologous protein sequences. </div> <div>Each eukaryotic cluster and all its significant prokaryotic hits constituted an individual</div> <div> sequence set, hereinafter referred to as an Eukaryotic/Prokaryotic Orthologous Cluster </div> <div>(EPOC). The EPOCs constitute groups of homologous proteins from eukaryotes and prokaryotes </div> <div>(each EPOC contains a unique set of eukaryotic proteins, but some clusters of prokaryotic </div> <div>proteins can be present in multiple EPOCs) that were used for phylogenetic tree </div> <div>construction, annotation, and evolutionary hypothesis testing. </div> <div> </div> <div>To infer the most likely prokaryotic ancestry of the eukaryotic proteins in each EPOC, </div> <div>rather than relying on the tree topology directly, we employed a probabilistic approach </div> <div>for evolutionary hypothesis testing using constraint trees. We exhaustively sampled all </div> <div>arrangements of likely sister clades and obtained Expected Likelihood Weights (ELW) for </div> <div>the set of possible sister clade models. As the ELW metric is analogous to model selection </div> <div>confidence, here we take it to be proportional to the probability of a sampled prokaryotic </div> <div>clade to be the true sister group of the given eukaryotic clade among a set of competing </div> <div>sister clades. For each EPOC, our analysis dynamically accounts for long branch outliers </div> <div>and is robust to phylogenetically non-homogenous clades. This analysis is further capable </div> <div>of resolving eukaryotic paraphyly, treating each eukaryotic clade within a EPOC as a </div> <div>single datapoint for downstream analysis. Our resulting data contains EPOCs annotated </div> <div>using profiles generated from KEGG Orthology Groups (KOGs), each with an MSA generated </div> <div>using muscle5, a maximum likelihood tree inferred using IQtree2 and associated ELW values </div> <div>for all candidate prokaryotic sister phyla. The analysis of prokaryotic ancestry was </div> <div>performed only for those eukaryotic clades that included more than 5 distinct taxonomic </div> <div>labels, with at least one coming from Amorphea and one from Diaphoretickes, the two </div> <div>expansive eukaryotic clades considered to represent either the first or the second </div> <div>bifurcation in the evolution of eukaryotes. Thus, these clades likely represent genes </div> <div>mapping back to the LECA.</div> <div> </div> <div>For further details please see main publication or contact</div> <div>victor.tobiasson@nih.gov</div> <div>eugene.koonin@nih.gov</div> <div> </div> <div> </div> <div>### Included files</div> <div> </div> <div>Unless otherwise stated all files contained are tab separated and utf-8 encoded </div> <div>with the first row containing header information. </div> <div>All data entries encoding lists are “|” (pipe) separated. </div> <div>Fields without data values are filled with string entries of “none”.</div> <div> </div> <div>--- Databases ---</div> <div>euk72_ep.tar.gz</div> <div>prok2311_as.tar.gz</div> <div>Prok2311As_final_clusters.tsv</div> <div>Euk72Ep_final_clusters.tsv</div> <div>prok2311_as.hmmDB.tar.gz</div> <div>euk72_ep.hmmDB.tar.gz</div> <div> </div> <div>--- Annotation and Curation ---</div> <div>NCBI_taxonomy_species_addendum.tsv</div> <div>NCBI_taxonomy_class_addendum.tsv</div> <div>Euk72Ep_Prok2311As_final_classes.tsv</div> <div>Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>KEGG_category_mapping.tsv</div> <div>KEGG_metadata.tsv</div> <div> </div> <div>--- EPOC data ---</div> <div>EPOC_data.tar.gz</div> <div>EPOC_annotation_KEGG.tsv</div> <div>EPOC_data.tsv</div> <div>EPOC_data.pangenomes_s10.tsv</div> <div>EPOC_data.pangenomes_s25.tsv</div> <div>EPOC_data.pangenomes_s67.tsv</div> <div>EPOC_data.GTDB.tsv</div> <div> </div> <div># euk72_ep.tar.gz</div> <div>Gunzip-ed .tar archive containing a single directory with 10 files </div> <div>constituting the initial eukaryotic mmseqs2 database with taxonomy annotation. </div> <div>Constructed from a pre-selected list of 72 eukaryotic proteomes downloaded from </div> <div>NCBI as well as a “clean” version of Eukprot, lacking highly prokaryotic-like </div> <div>contaminant sequences. </div> <div> </div> <div># prok2311_as.tar.gz</div> <div>Gunzip-ed .tar archive containing a single directory with 10 files constituting the </div> <div>initial prokaryotic mmseqs2 database with taxonomy annotation. Constructed from </div> <div>47545 complete genomes retrieved from NCBI in November 2023. </div> <div> </div> <div># prok2311_as.hmmDB.tar.gz</div> <div>Gunzip-ed .tar archive containing 6 files. Comprises an HHSuite Databse formatted </div> <div>from prok2311_as non--singleton clusters, contains 26286 profiles.</div> <div> </div> <div># euk72_ep.hmmDB.tar.gz</div> <div>Gunzip-ed .tar archive containing 6 files. Comprises an HHSuite Databse formatted </div> <div>from euk72_ep non-singleton clusters, contains 1631704 profiles.</div> <div> </div> <div># NCBI_taxonomy_species_addendum.tsv</div> <div>Taxonomy mapping file with manually curated ‘class’ level annotation for poorly </div> <div>annotated species. </div> <div> </div> <div>taxid: NCBI taxid</div> <div>proposed_class_id: Manually assigned NCBI taxid</div> <div>proposed_class_label: NCBI class name</div> <div>org_name: NCBI organism name</div> <div> </div> <div># NCBI_taxonomy_class_addendum.tsv</div> <div>Class revision file mapping poorly populated class level entries to higher order </div> <div>manually curated labels. Also includes information for small classes with shallow </div> <div>taxonomy which are deleted from the EPOC analysis at the level of tree construction.</div> <div> </div> <div>taxid: NCBI taxid</div> <div>ncbi_class: NCBI taxid of rank corresponding to ‘class’ following manual </div> <div>amendment as per NCBI_taxonomy_species_addendum.tsv</div> <div>revised_class_id: Manually assigned NCBI taxid of rank corresponding to ‘class’</div> <div>revised_class_label: Proposed cleartext name of manually revised revised_class_id </div> <div> </div> <div># Euk72Ep_Prok2311As_final_classes.tsv</div> <div>Final taxonomy at NCBI rank ‘class’ following revisions for all sequences in Euk72Ep or </div> <div>Prok2311As. These taxonomic labels are used for EPOC tree annotation. </div> <div> </div> <div>acc: mmseqs database header in either prok2311_as or euk72_ep databases</div> <div>taxid: NCBI taxid for organism</div> <div>superkingdom: Top level NCBI taxonomy classification Bacteria, Archaea or Eukarya, </div> <div>used to define Eukaryotic outgroups in EPOC analysis</div> <div>class: Cleartext name of manually revised NCBI rank ‘class’ identifier for annotation</div> <div> </div> <div># Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>Final taxonomy at GTDB rank ‘phylum’ transferred using marker genes from GTDB release 220</div> <div> </div> <div>acc: mmseqs database header in either prok2311_as or euk72_ep databases</div> <div>taxid: NCBI taxid for organism</div> <div>superkingdom: Top level NCBI taxonomy classification Bacteria, Archaea or Eukarya, </div> <div>used to define Eukaryotic outgroups in EPOC analysis</div> <div>class: Cleartext name of assigne GTDB phylum</div> <div> </div> <div># Prok2311As_final_clusters.tsv</div> <div>Cluster mapping file for accessions within the initial Prok2311A database to the </div> <div>final clusters used for HMM creation </div> <div> </div> <div>cluster_acc: cluster representative</div> <div>acc: cluster member</div> <div> </div> <div># Euk72Ep_final_clusters.tsv</div> <div>Cluster mapping file for accessions within the initial Prok2311A database to the </div> <div>final clusters used for HMM creation</div> <div> </div> <div>cluster_acc: cluster representative</div> <div>acc: cluster member</div> <div> </div> <div># EPOC_data.tar.gz</div> <div>Gunzip-ed directory containing 16035 EPOC folders. Each folder is named corresponding </div> <div>to the eukaryotic cluster representative which generated its profile as an ID </div> <div>Matches the tree_name field in EPOC_data_prok2311As.tsv</div> <div>contains the following files:</div> <div> </div> <div><EPOC_ID>.merged.fasta: sequences for all members of the EPOC</div> <div><EPOC_ID>.merged.fasta.leaf_mapping: tsv separated file containing taxonomy and tree reduction data</div> <div><EPOC_ID>.merged.fasta.muscle: main cropped MSA for tree generation </div> <div><EPOC_ID>.merged.fasta.muscle.iqtree: IQtree2 output from tree generation</div> <div><EPOC_ID>.merged.fasta.muscle.treefile.annot: annotated newick tree file with final tree</div> <div><EPOC_ID>.merged.tree_data.tsv: final parsed tree data with columns matching EPOC_data_prok2311As.tsv</div> <div> </div> <div>EPOCs with more than one possible eukaryotic sister phyla also contains </div> <div>a folder "constraint_analysis" with constraint tree information used for </div> <div>ELW value calculation. </div> <div> </div> <div># EPOC_data.tsv</div> <div>Main resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) </div> <div>based on pangenomes defined as including 10% of species per class. This is the main</div> <div>data to be used for genereting the core dataset and for data visualistation</div> <div>Contains information regarding tree breakdown, LCA membership and phylogenetic </div> <div>distances between all detected LCAs. Equivalent to the stacked dataframes from all </div> <div>EPOC directories in EPOC_data </div> <div> </div> <div>tree_name: unique index for each EPOC </div> <div>euk_clade_rep: unique index for each annotated eukaryotic clade within each tree_name</div> <div>euk_clade_size: number of original sequences represented by euk_clade_rep</div> <div>euk_clade_weight: metric for taxonomic purity for each euk_clade_rep</div> <div>euk_leaf_clade: boolean indicating whether euk_clade_rep contains a single leaf</div> <div>euk_LCA: lowest taxa spanning all members in euk_clade_rep</div> <div>euk_scope: list of all taxonomic classes in euk_clade_rep</div> <div>euk_scope_len: length of euk_scope list</div> <div>prok_clade_rep: unique index for each annotated prokaryotic clade for each euk_clade_rep</div> <div>prok_clade_size: number of original sequences represented by prok_clade_rep</div> <div>prok_clade_weight: metric for taxonomic purity for each prok_clade_rep</div> <div>prok_leaf_clade: boolean indicating whether prok_clade_rep contains a single leaf</div> <div>prok_taxa: lowest taxa spanning all members in prok_clade_rep</div> <div>dist: tree-distance from lowest tree node containing all members of prok_clade_rep to lowest tree node containing all members of euk_clade_rep</div> <div>top_dist: graph-distance (node-distance) from lowest tree node containing all members of prok_clade_rep to lowest tree node containing all members of euk_clade_rep</div> <div>raw_stem_length: tree-distance from lowest tree node containing the union of all members of prok_clade_rep and euk_clade_rep to the tree node containing all members of euk_clade_rep</div> <div>median_euk_leaf_dist: median value for all tree distances from the tree node containing all members of euk_clade_rep to the individual leaves</div> <div>stem_length: raw_stem_length/median_euk_leaf_dist</div> <div>logL: log likelihood of best constraint tree constructed</div> <div>deltaL: log likelihood difference between constraint tree for prok_clade_rep and best constraint tree constructed</div> <div>bp-RELL: validation metric from IQtree -trees, see iqtree.org</div> <div>bp-RELL_accept: as above</div> <div>p-KH: as above</div> <div>p-KH_accept: as above</div> <div>p-SH: as above</div> <div>p-SH_accept: as above</div> <div>c-ELW: as above</div> <div>c-ELW_accept: as above</div> <div>p-AU: as above</div> <div>p-AU_accept: as above</div> <div> </div> <div># EPOC_data.pangenomes_s10.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated </div> <div>based on pangenomes defined as including 10% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div> </div> <div># EPOC_data.pangenomes_s25.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated </div> <div>based on pangenomes defined as including 25% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div> </div> <div># EPOC_data.pangenomes_s67.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated </div> <div>based on pangenomes defined as including 67% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div> </div> <div># EPOC_data.GTDB.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated </div> <div>under revised taxonomy from GTDB based on data from Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>Identical file structure to EPOC_data.tsv</div> <div> </div> <div># EPOC_data.alpha_replicates.tsv</div> <div>Resulting data from 20 repetitions of Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated </div> <div>from a subset of Alphaproteobacterial-derived EPOCs. </div> <div>Identical file structure to EPOC_data.tsv with the addition of:</div> <div> </div> <div>rep: indicating technical replicate number, 0-19</div> <div> </div> <div># EPOC_annotation_KEGG.tsv</div> <div>Parsed HHblits output of HMM profiles generated from KEGG KOGs (KEGG Orthologous Groups) </div> <div>against eukaryotic profiles constituting each EPOC</div> <div> </div> <div>Query: query name equal to tree_name from EPOC_data</div> <div>Target: target name equal to kogid in KEGG_category_mapping and KEGG_metadata</div> <div>Prob: data from HHblits, see https://github.com/soedinglab/hh-suite/wiki</div> <div>E-value : as above</div> <div>P-value : as above</div> <div>Score: as above</div> <div>SS: as above</div> <div>Cols: as above</div> <div>Identities: as above</div> <div>Similarity: as above</div> <div>Sum_probs: as above</div> <div>Query-HMM-start: as above</div> <div>Query-HMM-end: as above</div> <div>Template-HMM-start: as above</div> <div>Template-HMM-end: as above</div> <div>Template_columns: as above</div> <div>Template_Neff : as above</div> <div>Pairwise_cov: calculated pairwise coverage from Query and Target start and end</div> <div>Description: category_name from KEGG_category_mapping</div> <div> </div> <div># KEGG_category_mapping.tsv</div> <div>Mapping of relevant KOG identifiers to their higher order categories as </div> <div>"Maps" "Modules" or "Reactions" as per KEGG see https://www.kegg.jp/kegg/pathway.html</div> <div> </div> <div>kogid: unique KOG identifier</div> <div>category_id: KEGG map, module, or reaction number</div> <div>category_name: cleartext name for KOG identifier</div> <div> </div> <div># KEGG_metadata.tsv</div> <div>File mapping KOGs to BRITE classification and to additional databases of chemical properties.</div> <div> </div> <div>kogid: unique KOG identifier</div> <div>name: cleartext name for KOG identifier</div> <div>brite_A: list of BRITE-A sets including KOG</div> <div>brite_B: list of BRITE-A sets including KOG</div> <div>brite_C: list of BRITE-A sets including KOG</div> <div>EC: list of Enzyme commission numbers associated with KOG, see https://enzyme.expasy.org/</div> <div>TC: list of transporter classification numbers associated with KOG, see https://www.tcdb.org/</div> <div>RN: list of KEGG reaction numbers associated with KOG</div> <div>CA: list of CAZY numbers associated with KOG, see http://www.cazy.org/</div> <div>GO: list of GO terms associated with KOG, see https://geneontology.org/</div> </div> <div> </div> </div>
Intermediate processing stage of horizontal particle flux data collected using a snow particle counter on board the R/V Akademik Tryoshnikov in the Southern Ocean during the austral summer of 2016/17 as part of the Antarctic Circumnavigation Expedition (ACE).
<p><strong>Dataset abstract</strong></p> <p>Flux of particles (snow, rain and other particles including sea spray) were recorded passing through a photo-electric snow particle counter installed on board the R/V Akademik Tryoshnikov as part of the Antarctic Circumnavigation Expedition (ACE). Data were recorded from January to March 2017 in the Southern Ocean. Here we present an intermediate step in data processing, with relative horizontal particle flux of particles with a size between 36 – 2000 μm averaged over one-minute periods. Data are presented in daily files.</p> <p><strong>Dataset contents</strong></p> <ul> <li>SPC_HPF_1min_YYYY_MM_DD.csv, data files, comma-separated values</li> <li>data_file_header.txt, metadata, text</li> <li>README.txt, metadata, text</li> </ul> <p><strong>Dataset license</strong></p> <p>This one-minute averaged horizontal particle flux dataset from ACE is made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0) whose full text can be found at https://creativecommons.org/licenses/by/4.0/</p>
GIXD data of organic-inorganic methylammonium lead bromide perovskite (MAPbBr3), processed q-space maps
<p>This dataset contains grazing incidence x-ray diffraction (GIXD) maps projected in q-space and polar projection. The underlying raw data is published in <a href="https://doi.org/10.5281/zenodo.6683616">10.5281/zenodo.6683616</a> and processed with <a href="https://doi.org/10.5281/zenodo.6683658">10.5281/zenodo.6683658</a>. This data describes a time series of diffraction images acquired with 10 Hz.</p> <p> </p> <p>Parameters of the provided data:</p> <ul> <li> <p>Q-space-maps</p> </li> </ul> <p> </p> <ul> <li> <ul> <li> <p>Horizontal axis (Q<sub>xy</sub>) range: (0, 3.2) Å<sup>-1</sup></p> </li> <li> <p>Vertical axis (Q<sub>z</sub>) range: (0, 3.2) Å<sup>-1</sup></p> </li> <li> <p>Resolution: 1350x1350 pixels</p> </li> <li> <p>Origin (lower left coordinate in q): (0, 0)</p> </li> </ul> </li> <li> <p>Polar data</p> <ul> <li> <p>Horizontal axis (||<strong>q</strong>||) range: (0, 4.53) Å<sup>-1</sup></p> </li> <li> <p>Vertical axis (ф) range: (0, 90) deg</p> </li> <li> <p>Resolution: 512x1024 pixels</p> </li> <li> <p>Origin (lower left coordinate in q): (0, 0)</p> </li> </ul> </li> </ul>
In-situ grazing-incidence X-ray diffraction data of the crystallization process of organic-inorganic methylammonium lead bromide perovskite (MAPbBr3) via employing an isopropanol antisolvent. Raw Data
<p>The dataset contains 400 diffraction images from a 40 second in-situ grazing-incidence wide-angle X-ray scattering measurement of the crystallization process of organic-inorganic methylammonium lead bromide perovskite (MAPbBr3) on a glass substrate. The crystallization is initiated via employing an isopropanol antisolvent during the spin-coating of the perovskite precursor solution. 40 µL of MAPbBr3 solution (4:1 DMF/DMSO solvent mixture) was applied on plasma-cleaned glass substrate in a chamber with kapton windows. The two-phase spin-coating regime included 10 seconds at 1000 rpm followed by 30 seconds at 2000 rpm, 200 µL of antisolvent was dispensed at t = 30 s.</p> <p> </p> <p> </p> <p>The data was acquired at the P08 Beamline at PETRA III (DESY Hamburg). Acquisition parameters:</p> <p> </p> <ul> <li> <p>X-ray wavelength: 0.6888 nm</p> </li> <li> <p>Sample detector distance: 809 mm</p> </li> <li> <p>Incidence angle: 0.5 deg.</p> </li> <li> <p>Detector model: XRD 1621 CN3 EHS</p> </li> <li> <p>Acquisition rate : 10 frames per second (10 Hz)</p> </li> <li> <p>Direct beam position (pixels): 545, 222</p> </li> </ul>
Processed glider data: 9 months of hydrographic and ADCP observations in the Gulf of Oman.
<p>68 repeat transects and 2 virtual moorings covering a spring/neap cycle collected by a SeaExplorer glider with T, S, O2, Chl, Optical backscatter, PAR and ADCP data in the Gulf of Oman. Dataset collected as part of the ONR Global project "Shelf slope dyanmics in the Sea of Oman: How submesoscale processes control food and water security". The glider was deployed from the north shore of Oman into the Gulf of Oman, sampling down to 1000m in the oxygen minimum zone.</p> <p> </p> <p>File and variable metadata included in the netCDF files.</p> <p> </p> <p>sea057_M##.ad2cp.#####.nc : Raw ADCP data provided in Nortek .nc format. (version 1.0)</p> <p>SEA057_glider.nc : SeaExplorer data timeseries QC'd and processed into a 1Hz timeseries. (version 1.0)</p> <p>SEA057_ADCP_v2.nc : ADCP data fully processed, binned (2 dbar) and referenced, and then reprojected back onto a timeseries. ADCP data processed as per https://github.com/bastienqueste/gliderad2cp . (version v3)</p> <p> </p>
Labeled Time Series Data of Force/Torque for Monitoring Assembly Processes with a Delta Robot
<p>This dataset comprises 524 recordings of 6-dimensional time series data, capturing forces in three directions and torques in three directions during the assembly of small car model wheels. The data was collected using an equidistant sampling method with a sampling period of 0.004 seconds. Each time series represents the process of assembling one wheel, specifically the placement of a tire onto a rim, and includes a label indicating whether the assembly was successful (OK). The wheels were assembled in batches of four, and the recordings were obtained over six different days. The labels of recordings from two (days 3 and 4) of the six days are invalid as described in [1]. The labels presented in this data set are only binary (they do not describe the reason of the failure). The labels of recordings from days 5 and 6 are created by human while the other labels came from a convolutional neural network based computer vision classifier and can be inaccurate as described in section 5.4 of [1]. </p> <h4>Dataset Structure:</h4> <ul> <li><strong>File:</strong> <code>ForceTorqueTimeSeries.csv</code> <ul> <li><strong>Columns:</strong> <ul> <li><code>idx (1-524)</code>: Index of the recording corresponding to the assembly of one wheel.</li> <li><code>label (true/false)</code>: Indicates whether the assembly was successful (TRUE = product is OK).</li> <li><code>meas_id (1-6)</code>: Identifier for the day on which the recording was made (refer to Table 2.1 in [1]).</li> <li><code>force_x</code>: X-component of the force measured by the sensor mounted on the delta robot's end effector.</li> <li><code>force_y</code>: Y-component of the force.</li> <li><code>force_z</code>: Z-component of the force.</li> <li><code>torque_x</code>: X-component of the torque.</li> <li><code>torque_y</code>: Y-component of the torque.</li> <li><code>torque_z</code>: Z-component of the torque.</li> </ul> </li> </ul> </li> </ul> <h4>Additional Files:</h4> <ul> <li><strong><code>IMG_3351.MOV</code>:</strong> A video demonstrating the assembly process for one batch of four wheels.</li> <li><strong><code>F3-BP-2024-Trna-Ales-Ales Trna - 2024 - Anomaly detection in robotic assembly process using force and torque sensors.pdf</code>:</strong> Bachelor thesis [1] detailing the dataset and preliminary experiments on fault detection.</li> <li><strong><code>F3-BP-2024-Hanzlik-Vojtech-Anomaly_Detection_Bachelors_Thesis.pdf</code>:</strong> Bachelor thesis [2] describing the data acquisition process.</li> </ul> <h3>References:</h3> <ol> <li>Trna, A. (2024). <em>Anomaly detection in robotic assembly process using force and torque sensors</em> [Bachelor’s thesis, Czech Technical University in Prague].</li> <li>Hanzlik, V. (2024). <em>Edge AI integration for anomaly detection in assembly using Delta robot</em> [Bachelor’s thesis, Czech Technical University in Prague].</li> </ol>
EISCAT Svalbard radar Common Program data from February 26 to February 28 2023, which is processed by GUISDAP
<p>This is two-dimensional (time and altitude) ionospheric parameter data that contains electron density, electron temperature and ion temperature. It is estimated based on EISCAT Svalbard radar measurement implemeted as common program from February 26 to Feburuary 28, 2023 (https://portal.eiscat.se/) and processed by a software for incoherent scatter radar analysis, GUISDAP (https://gitlab.com/eiscat/guisdap9). The more detailed descriptions can be found as metadata in the uploaded netCDF file.</p>
Data for: Techno-economic analysis of a novel laccase production process utilizing perennial biomass and the aqueous phase of bio-oil, Iowa, USA 2023-2025
This dataset contains the experimental design, measurements, and derived variables used to parameterize a techno‑economic model of laccase production via two‑stage solid‑state fermentation of prairie biomass with bio‑oil aqueous phase induction. It includes nutrient screening data for Pleurotus ostreatus growth on prairie biomass with alternative nitrogen sources and a corn‑steep solids dose series; factorial/response‑surface experiments varying substrate bed depth, substrate‑to‑inoculum (S:I) ratio, and pre‑induction growth time; and time‑resolved induction measurements. For each run and replicate, the data record the full set of spectrophotometric absorbances at 0–210 s, fitted slopes and r-square values, dilution and volume factors, and laccase activities normalized per mL and per gram of biomass, alongside the exact culture timings and environmental conditions used in the ABTS assay at 420 nm. Results tables provide the fitted central‑composite design model terms (coefficients, F‑statistics, and p‑values) used directly as inputs to the minimum laccase selling price (MLSP) calculations, together with the underlying per‑condition raw results.
Larval transport pathways from three prominent sand lance habitats in the Gulf of Maine: otolith data, model data, and post-processed model data products
This dataset includes hatch and larval period for sand lance collected in 2019 and results from particle tracking runs of simulated sand lance larvae throughout the Northeast U.S. Shelf as part of Long-Term Ecological Research (NES-LTER). Release dates vary by region, corresponding to hatch and settlement dates of settling sand lance collected in 2019. Particles were depth-keeping throughout the upper 40 m to best replicate our understanding of the vertical distribution of sand lance larvae. Data were used to determine the average particle transport pathways from these sand lance habitats, including connectivity among the three hotspots, and spatial variability of connectivity within each hotspot. Further information can be found within the manuscript: Suca, J. J., Ji, R., Baumann, H., Pham, K., Silva, T. L., Wiley, D. N., Feng, Z., & Llopiz, J. K. (2022). Larval transport pathways from three prominent sand lance habitats in the Gulf of Maine. Fisheries Oceanography, 31( 3), 333-352. https://doi.org/10.1111/fog.12580
Effects of Multiple Resource Additions on Community and Ecosystem Processes: NutNet Seasonal Biomass and Seasonal and Annual NPP Data at the Sevilleta National Wildlife Refuge, New Mexico
Two of the most pervasive human impacts on ecosystems are alteration of global nutrient budgets and changes in the abundance and identity of consumers. Fossil fuel combustion and agricultural fertilization have doubled and quintupled, respectively, global pools of nitrogen and phosphorus relative to pre-industrial levels. In spite of the global impacts of these human activities, there have been no globally coordinated experiments to quantify the general impacts on ecological systems. This experiment seeks to determine how nutrient availability controls plant biomass, diversity, and species composition in a desert grassland. This has important implications for understanding how future atmospheric deposition of nutrients (N, S, Ca, K) might affect community and ecosystem-level responses. This study is part of a larger coordinated research network that includes more than 40 grassland sites around the world. By using a standardized experimental setup that is consistent across all study sites, we are addressing the questions of whether diversity and productivity are co-limited by multiple nutrients and if so, whether these trends are predictable on a global scale. Above-ground net primary production is the change in plant biomass, represented by stems, flowers, fruit and and foliage, over time and incoporates growth as well as loss to death and decomposition. To measure this change the vegetation variables, including species composition and the cover and height of individuals, are sampled twice yearly (spring and fall) at permanent 1m x 1m plots within each site. Volumetric measurements are made using vegetation data from permanent plots (SEV231, "Effects of Multiple Resource Additions on Community and Ecosystem Processes: NutNet NPP Quadrat Sampling") and regressions correlating species biomass and volume constructed using seasonal harvest weights from SEV157, "Net Primary Productivity (NPP) Weight Data."
200 kHz pre-processed echosounder data collected on the Antarctic Circumnavigation Expedition during the austral summer of 2016/2017.
<p><strong>Dataset abstract </strong></p> <p>These data consist of pre-processed echosounder observations in the Southern Ocean collected during the Antarctic Circumnavigation Expedition (ACE; Leg2-Leg3) using an EK60 GPT operating at 200 kHz. The instrument was calibrated at South Georgia during the expedition (Leg 3) and corrections were applied prior to calculation of the volume backscattering strength (Sv). The signal-to-noise ratio (SNR) was analysed and was deemed very poor at depths greater than 100 m. Therefore, only data collected between the transducer depth (8.4 m) and 100 m were archived.</p> <p><strong>Dataset contents</strong></p> <ul> <li>ACE-DYYYYMMDD-THHMMSS.csv, data files, comma-separated values</li> <li>data_file_header.txt, metadata, text</li> <li>README.txt, metadata, text</li> </ul> <p><strong>Dataset license</strong></p> <p>This 200 kHz pre-processed echosounder data collected on ACE is made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0) whose full text can be found at https://creativecommons.org/licenses/by/4.0/</p>
IPBES Data Management Tutorials - Session 5.4: Processing and analysis
<p>The <em>IPBES data management tutorials</em> are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The<em> Tools for data management </em>chapter provides IPBES authors with an overview of open source tools used frequently by the scientific community to help it implement data management for the entire data life cycle.</p> <p>This session on <em>processing and analysis </em>reviews common scripting languages for data analysis and processing, such as python and R. </p>
OSeaIce post-processed data
<p>These are the EC-Earth3 post-processed data that can be used with the associated Python scripts (https://doi.org/10.5281/zenodo.4291971) in order to produce the figures of the paper that summarizes the results arising from the sensitivity experiments carried out in the framework of the EU Horizon 2020 OSeaIce project (Marie Slodowska-Curie project, grant agreement no. 834493):</p> <p>Docquier et al. (2020). Impact of ocean heat transport on the Arctic sea-ice decline: A model study with EC-Earth3. Accepted in Climate Dynamics.</p>
CompTox Zebrafish developmental toxicity processed data
<p>Please cite the published paper and original data source.</p> <p>A classification data set for 19 endpoints of Zebrafish developmental toxicity for QSAR modelling.</p> <p>The original data was downloaded from <a href="https://comptox.epa.gov/dashboard">https://comptox.epa.gov/dashboard</a></p> <p>See also references below.</p> <ul> <li>20201230_tx_zf.csv -- Endpoint data</li> <li>20201230_tx_zf_smiles.txt -- SMILES of the structures</li> <li>20201230_tx_zf_descriptors_rdkit.csv -- Descriptors</li> <li>20201230_tx_zf_fp_r3_v5120.csv -- Fingerprints</li> </ul>
Numerical weather simulation using COSMOiso in June 2019 during L-WAIVE field campaign: selected model output and post-processed data.
<p>This dataset consists of extracts from a simulation with the isotope-enabled regional numerical weather prediction model COSMOiso, which covers the timespan of the Lacustrine-Water vApor Isotope inVentory Experiment (L-WAIVE) field campaign taking place in June 2019 in the Annecy valley in the French Alps (Chazette et al. 2021).The simulation has a horizontal resolution of 0.1° (~10km) and 40 vertical levels.</p><p>This COSMOiso simulation is used in Thurnherr et al. (submitted) to compare stable water isotope measurements from various platforms. Here, we provide selected model outputs and post-processed data used in this comparison study. The post-processed data contain:</p><ol><li>COSMOiso output files for time steps 20190612_12, 20190613_12, 20190615_13, 20190616_13, 20190617_12, 20190622_12.</li><li>Pressure weighted total and subcolumn averages for time steps 20190612_12, 20190613_12, 20190615_13, 20190616_13, 20190617_12, 20190622_12.</li><li>Vertical cross section of selected variables at Annecy, the location of the L-WAIVE field campaign, for the simulation time window.</li><li>Interpolated time series of subcolumn and total column averages at Annecy, the location of the L-WAIVE field campaign, for the simulation time window.</li><li>Interpolated variables along the flight tracks from the L-WAIVE campaign (see Sodemann and Seidl, 2023).</li></ol><p>See also README files for more details on the provided data.</p><p>To access further model output and post-processed data, please contact the dataset authors.</p>
Public Available Data Set of Process Flows from Internal Physical Inspections in the Failure Analysis Laboratory
<p>This data set was generated in accordance with the semiconductor industry and contains data of certain process flows in Failure Analysis (FA) laboratories focusing on the identification and analysis of anomalies or malfunctions in semiconductor devices. It comprises logistic data about the processing steps for the so-called Internal Physical Inspection (IPI).</p><p>A so-called IPI job is given as a sequence of tasks that must be performed to complete the job they belong to. It has an assigned unique ID and timestamps indicating the submission, the end, and the deadline to be met. A job also has an IPI classification assigned to it, providing general guidelines on the operations to be performed.</p><p>Every task within a job has its own type and working time, as well as the assigned resources. There are two main resources involved:</p><p> - the equipment; the machine used to perform the task,</p><p> - the operator; the person who performed the task.</p><p>In addition, general information about the type of the device to be analyzed is also available, such as the given (anonymized) package and basictype. Data also include the number of stressed samples within a device and the samples a task is performed on.</p><p>The dataset includes data from 4 years, specifically from January 2020 to December 2022.</p><p>Finally, the exact column structure is given as follows (python 3.9.5 datatype):</p><ul><li>JOB_ID [int64]: the unique ID of the job</li><li>JOB_SUBMISSION_DATE [object]: the date of the job submission</li><li>JOB_REQ_END_DATE [object]: the required end date (deadline)</li><li>JOB_FINISH_DATE [object]: the actual end date</li><li>JOB_BASICTYPE_H [object]: the given basictype denotation</li><li>JOB_PACKAGE_H [object]: the package denotation of the device</li><li>JSH_QTY_STRESSED [float64]: number of stressed samples</li><li>TASK_SUBMISSION_DATE [object]: the date of the task submission</li><li>TASK_WORKING_TIME [float64]: the amount of time (hours) the task needs to be completed</li><li>TASK_SAMPLE_NO [object]: the samples the task was performed on </li><li>TASK_CEQ_ID [float64]: the ID of the machine used to perform the task</li><li>TASK_CTKS_ID [int64]: the ID representing the task type</li><li>TASK_USR_ID [int64]: the ID of the operator performing the task</li><li>CIPI_LEVEL_0 [object]: a series of IPI classifications, indicating what is required to execute for a specific job</li></ul>
3D Data Derivatives of Grotta di Fumane: GigaMesh-processed, Annotations and Segmentations
<p><strong>Overview:</strong></p> <p>This repository contains derivatives of the Open Access publication by Falcucci & Peresani [FP22]. Our derived dataset (n = 62) is used to demonstrate our segmentation algorithm [BHM23], as shown in [BLM22], [BLM23], [LBM23], and will serve as a benchmark dataset for future analyses. To date, and to the best of our knowledge, our dataset is the first dataset of annotated lithic artifacts. In addition to the annotated dataset, we will also provide the segmented [BLM23] and GigaMesh preprocessed datasets [Mar+10; MK13] (n = 732) in separate folders. </p> <p><strong>Repository description: </strong></p> <p>A detailed description of the data can be found in 3D_Data_Derivatives_of_GdF_overview.pdf.</p> <p>For information on the archaeological interpretation of the artifacts, please refer to the original data publication by Falcucci and Peresani (2022). In our publications, we have expanded the CSV file from Falcucci and Peresani (2022) to document the use of the extended dataset:</p> <ul> <li> <p>Annotated: All artifacts that are annotated are marked with a 1.</p> </li> <li> <p>GT_PLY: All artifacts that are annotated and included in this publication are referenced by their respective file, such as 31_gt_labels.ply.</p> </li> <li> <p>Bullenkamp_et_al_2022: Artifacts utilized in [BLM22] are marked with a 1 .</p> </li> <li> <p>Bullenkamp_et_al_2023: Artifacts utilized in [BLM23] are marked with a 1.</p> </li> <li> <p>Linsel_et_al_2023: Artifacts utilized in [LBM23] are marked with a 1.</p> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.