Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6,467

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

6,467 results for “processes”

Learn how ShareScore rates datasets ↗
edi60/100

North Temperate Lakes LTER Processed eddy covariance time series fluxes from tower located on roof of the CFL building oriented toward Lake Mendota 2012 - current

We calculated eddy covariance based fluxes of CO2, H2O, heat, and momentum to study lake-atmosphere exchanges since 2012. These data were collected by Ankur Desai from 2012 to present using a CSAT-3 sonic anemometer and LI-7500 gas analyzer located on the roof of the CFL building. A footprint model (Kljun) was used to screen for lake only data.

openCC (other)Dec 2022View details →
edi60/100

Cascade Project at North Temperate Lakes LTER Core Data Process Data 1984 - 2016

Data useful for calculating and evaluating primary production processes were collected from 6 lakes from 1984-2016. Chlorophyll a and pheophytin were measured by the same fluorometric method from 1984-2016. In some years chlorophyll and pheophytin were separated into size fractions (total, and a 'small' fraction that passed a 35 um mesh screen). Primary production was measured by the 14C method from 1984-1998. Dissolved inorganic carbon for primary production calculation was calculated from Gran alkalinity titration and air-equilibrated pH until 1987 when this method was replaced by gas chromatography. Until 1995 alkaline phosphatase activity was measured as an indicator of phosphorus deficiency.

openCC (other)Dec 2022View details →
zenodo56/100

Conserved regulation of RNA processing in somatic cell reprogramming

<p><strong>Data set 1. Transcript expression across human RNA-Seq samples: estimated read counts. </strong>The file contains estimated read counts, generated by kallisto (<a href="https://pachterlab.github.io/kallisto/">https://pachterlab.github.io/kallisto/</a>), for human transcripts and RNA-Seq samples used in this study (see Additional file 2 of the accompanying publication). The format is a compressed (GZIP) tab-separated transcript-by-sample matrix. Ensembl transcript identifiers and a combined Sequence Read Archive study/sample name identifier serve as row and column names, respectively.</p> <p><strong>Data set 2. Transcript expression across murine RNA-Seq samples: estimated read counts. </strong>As in Data set 1, but for mouse transcripts.</p> <p><strong>Data set 3. Transcript expression across simian RNA-Seq samples: estimated read counts. </strong>As in Data set 1, but for chimpanzee transcripts.</p> <p><strong>Data set 4. Transcript expression across across human RNA-Seq samples: estimated transcript abundances. </strong>As in Data set 1, but instead of read counts, transcript abundances in transcripts per million (TPM), as estimated by kallisto (<a href="https://pachterlab.github.io/kallisto/">https://pachterlab.github.io/kallisto/</a>), are listed. Format, column and row names as in Data set 1.</p> <p><strong>Data set 5. Transcript expression across murine RNA-Seq samples: estimated transcript abundances. </strong>As in Data set 4, but for mouse transcripts.</p> <p><strong>Data set 6. Transcript expression across simian RNA-Seq samples: estimated transcript abundances. </strong>As in Data set 4, but for chimpanzee transcripts.</p> <p><strong>Data set 7. Differential expression analyses across human RNA-Seq sample groups: log fold changes. </strong>The file contains log fold changes, inferred by edgeR (<a href="http://bioconductor.org/packages/release/bioc/html/edgeR.html">http://bioconductor.org/packages/release/bioc/html/edgeR.html</a>), for human genes and the RNA-Seq sample group contrasts listed in Additional file 3 of the accompanying publication in a compressed (GZIP) TSV gene-by-comparison matrix. Ensembl gene identifiers and a descriptive contrast identifier serve as row and column names, respectively.</p> <p><strong>Data set 8. Differential expression analyses across murine RNA-Seq sample groups: log fold changes. </strong>As in Data set 7, but for mouse genes.</p> <p><strong>Data set 9. Differential expression analyses across simian RNA-Seq sample groups: log fold changes. </strong>As in Data set 7, but for chimpanzee genes.</p> <p><strong>Data set 10. Differential expression analyses across human RNA-Seq sample groups: false discovery rates. </strong>The file contains false discovery rates (FDR) for the differential expression analyses summarized in Data set 7. Format, column and row names as in Data set 7.</p> <p><strong>Data set 11. Differential expression analyses across murine RNA-Seq sample groups: false discovery rates. </strong>As in Data set 10, but for mouse genes.</p> <p><strong>Data set 12. Differential expression analyses across simian RNA-Seq sample groups: false discovery rates. </strong>As in Data set 10, but for chimpanzee genes.</p> <p><strong>Data set 13. Quantification of alternative splicing events across human RNA-Seq samples. </strong>The file contains &lsquo;percent spliced in&rsquo; (PSI) values computed by SUPPA (<a href="https://github.com/comprna/SUPPA">https://github.com/comprna/SUPPA</a>) for annotated alternative splicing events (inferred from the transcript annotation of the human genome, Ensembl release 84; <a href="http://www.ensembl.org/">http://www.ensembl.org/</a>). The format is a compressed (GZIP) tab-separated transcript-by-sample matrix. SUPPA-provided event identifiers and a combined Sequence Read Archive study/sample name identifier serve as row and column names, respectively.</p> <p><strong>Data set 14. Quantification of alternative splicing events across murine RNA-Seq samples. </strong>As in Data set 13, but for mouse alternative splicing events.</p> <p><strong>Data set 15. Differential splicing analyses across human RNA-Seq sample groups: differences in &lsquo;percent spliced in&rsquo; (&Delta;PSI). </strong>The file contains &Delta;PSI values for human alternative splicing events (as in Data set 13). The RNA-Seq sample group contrasts are listed in Additional file 3 of the accompanying publication. Values were inferred by SUPPA&rsquo;s diffSplice functionality (<a href="https://github.com/comprna/SUPPA">https://github.com/comprna/SUPPA</a>). The format is a compressed (GZIP) tab-separated gene-by-comparison matrix. SUPPA event identifiers and a descriptive contrast identifier serve as row and column names, respectively.</p> <p><strong>Data set 16. Differential splicing analyses across murine RNA-Seq sample groups: differences in &lsquo;percent spliced in&rsquo; (&Delta;PSI). </strong>As in Data set 15, but for mouse alternative splicing events.</p> <p><strong>Data set 17. Differential splicing analyses across human RNA-Seq sample groups: P values. </strong>The file contains P values for the differential splicing analysis of human alternative splicing events summarized in Data set 15. Format, column and row names as in Data set 15.</p> <p><strong>Data set 18. Differential splicing analyses across murine RNA-Seq sample groups: P values. </strong>The file contains P values for the differential splicing analysis of mouse alternative splicing events summarized in Data set 16. Format, column and row names as in Data set 15.</p> <p><strong>Data set 19. Transcript expression across murine RNA-Seq time course data: estimated read counts. </strong>As in Data set 2, but for the time course data generated for the accompanying publication.</p> <p><strong>Data set 20. Transcript expression across murine RNA-Seq time course data: estimated transcript abundances. </strong>As in Data set 5, but for the time course data generated for the accompanying publication.</p> <p><strong>Data set 21. Quantification of alternative splicing events across murine RNA-Seq time course data. </strong>As in Data set 14, but for the time course data generated for the accompanying publication.</p>

opencc-by-4.0Mar 2018View details →
zenodo56/100

InSAR stack of Fernandina volcano in Galápagos, Ecuador from Sentinel-1 descending track 128 processed with ISCE2/topsStack

<p>A stack of unwrapped interferograms on Fernandina volcano, Gal&aacute;pagos, Ecuador</p> <p>Sensor: Sentinel-1descending track 128</p> <p>Processor: ISCE/topsStack</p> <p>Tropospheric delay estimated from ERA-5&nbsp;using PyAPS is attached.</p> <p>This is an input dataset for the time series analysis with&nbsp;<a href="https://github.com/insarlab/MintPy/">MintPy</a>.</p> <p><strong>Version 1.x (~750 MB)</strong><br> Time: 2014.12.13 - 2018.06.19&nbsp;(98 acquisitions, 288 interferograms)</p> <p><strong>Version 0.1&nbsp;(~280 MB; for fast testing of code development)</strong><br> Time: 2014.12.13 - 2016.05..24 (36 acquisitions, 102 interferograms)</p>

opencc-by-4.0Feb 2019View details →
edi56/100

Processed Net N2 Flux and Nitrous Oxide Production Rates from a mesocosm experiment, North River, MA, 2022

The data documented here show nitrogen cycling rates (net N2 flux and net N2O flux) from a tidal, freshwater wetland under two stressors in isolation and in combination over two different disturbance regimes. We used intact core mesocosms to examine how nitrogen cycling changed in response to increased temperature and salinity under pulse and press disturbances. We found that net N2 flux rates, defined as the balance between nitrogen fixation and denitrification did not directionally change in response to stressor pulse or press. Instead, it became more variable under both disturbance regimes. Nitrous oxide production rates, however, decreased and became more stable over time in the press scenario, but remained highly variable in the pulse scenario. These findings provide valuable knowledge on the functional potential of the nitrogen cycling microbial communities in tidal, freshwater wetlands when facing future climate variability.

openCC0Oct 2025View details →
edi56/100

MCR LTER: Coral Reef: Dead coral skeletons impair key recovery processes following coral bleaching; data for Kopecky et al., 2024 Global Change Biology

The data included in this data package were collected on the North shore of Moorea, French Polynesia, from 2015-2023 to explore how dead coral skeletons (e.g,, left after coral bleaching events) influence critical processes tied to coral reef resilience. Together, these various datasets were used for analyses in the manuscript entitled "Changing disturbance regimes, material legacies, and stabilizing feedbacks: dead coral skeletons impair key recovery processes following coral bleaching", published in Global Change Biology. These data are in support of a publication Kopecky et al. (2024) Global Change Biology, and were a part of the thesis of K. Kopecky. The manuscript title and author list are as follows: Changing disturbance regimes, material legacies, and stabilizing feedbacks: dead coral skeletons impair key recovery processes following coral bleaching. Kai Kopecky, Russell J. Schmitt, Sally J. Holbrook. This material is based upon work supported by the U.S. National Science Foundation under Grant No. OCE 22-24354 (and earlier awards) as well as a generous gift from the Gordon and Betty Moore Foundation. Research was completed under permits issued by the French Polynesian Government (Délégation à la Recherche) and the Haut-commissariat de la République en Polynésie Francaise (DTRT) (Protocole d'Accueil 2005-2024). This work represents a contribution of the Moorea Coral Reef (MCR) LTER Site.

openCC (other)Aug 2024View details →
edi56/100

Cascade Project at North Temperate Lakes LTER: Process Data 1984 - 2007

Data on chlorophyll, primary productivity, and alkaline phosphatase activity from 1984-95. Samples were collected with a Van Dorn bottle at 6 depths determined from the percent of surface irradiance (100%, 50%, 25%, 10%, 5% and 1%) and in the hypolimnion (12 m in Peter, East Long, West Long, and Tuesday lakes; 9 m in Paul Lake; and 4.5 m in Central Long Lake). Sampling Frequency: varies Number of sites: 8

openCC (other)Nov 2022View details →
edi56/100

Fertilization (NPK) alter dryland biogeochemical processes

Nutrient augmentation is one major global change disturbance that could have cascading effects on local plant and microbial communities thus altering biogeochemical properties (Peñuelas et al. 2012). While many studies have investigated fertilization effects on community change and ecosystem processes, less work has been done in dryland ecosystems (Schimel 2010), where nutrient availability often comes as pulses correlated with rain events (Collins et al. 2008). We leveraged an ongoing fertilization experiment (NutNet) at the Sevilleta to answer the question: How does fertilization alter dryland biogeochemical processes, and how does this effect change seasonally? To explore this topic, we specifically measure three important soil hydrolase enzymes, N-acetyl- glycosaminidase (NAG), phosphatase (AP), and β- glucosidase (BG), microbial biomass, and soil nitrogen levels at 5 points along a seasonal gradient within the NutNet plots.

openCC (other)Apr 2022View details →
OpenNeuro52/100

Social Processes Initiative in Neurobiology of the Schizophrenia(s) Traveling Human Phantoms

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
OpenNeuro52/100

Brain Correlates of Math Processing in Adults

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
OpenNeuro52/100

Component processes of word reading in adults and children

Open the record for dataset details and reuse information.

openCC0Jan 2021View details →
zenodo52/100

Core collapse supernova yield from the post-processing of a long-term 3D simulation

<p>This dataset accompanies the publication<i> "Production of 44Ti and Iron-group Nuclei in the Ejecta of 3D Neutrino-driven Supernovae"</i> published in the <i>Astrophysical Journal Letters</i> Volume <strong>957</strong>, Issue 2, id.L25.</p><p>The dataset consists of an ACII text file that contains the isotopic yields from the post-processing of a 3D long-term supernova simulation for a 18.88 solar mass progenitor model. The yields are given in units of solar masses.&nbsp;</p><p><strong>Important: The dataset does not include the full stellar yield. </strong>It only represents the inner 0.142 solar masses. The total ejecta mass is expected to be larger.&nbsp;</p><p>The dataset is also available on the websites of the Max-Planck Institute for Astrophysics in Garching, Germany: https://wwwmpa.mpa-garching.mpg.de/ccsnarchive/data/Sieverding2023/</p><p>The results have been obtained using the open source nuclear reaction network code <a href="https://github.com/starkiller-astro/XNet">XNet.</a></p><p>Calculations have been performed on the supercomputing cluster Cobra the Max-Planck Computing and Data Facility (MPCDF) in Garching, Germany.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo52/100

Pre-processed (in Detectron2 and YOLO format) planetary images and boulder labels collected during the BOULDERING Marie Skłodowska-Curie Global fellowship

<p>This database contains 4976 planetary images of boulder fields located on Earth, Mars and Moon. The data was collected during the BOULDERING Marie Skłodowska-Curie Global fellowship between October 2021 and 2024. The data was already splitted into train, validation and test datasets, but feel free to re-organize the labels at your convenience.&nbsp;</p> <p>For each image, all of the boulder outlines within the image were carefully mapped in QGIS. More information about the labelling procedure can be found in the following manuscript (<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013</a>). This dataset differs from the previous dataset included along with the manuscript&nbsp;<a href="https://zenodo.org/records/8171052">https://zenodo.org/records/8171052</a>, as it contains more mapped images, especially of boulder populations around young impact structures on the Moon (cold spots). In addition, the boulder outlines were also pre-processed so that it can be ingested directly in YOLOv8.</p> <p>A description of what is what is given in the README.txt file (in addition in how to load the custom datasets in Detectron2 and YOLO). Most of the other files are mostly self-explanatory. Please see previous dataset or manuscript for more information. If you want to have more information about specific lunar and martian planetary images, the IDs of the images are still available in the name of the file. Use this ID to find more information (e.g., M121118602_00875_image.png, ID M121118602 ca be used on https://pilot.wr.usgs.gov/). I will also upload the raw data from which this pre-processed dataset was generated (see <a href="https://zenodo.org/records/14250970">https://zenodo.org/records/14250970</a>).</p> <p>Thanks to this database, you can easily train a Detectron2 Mask R-CNN or YOLO instance segmentation models to automatically detect boulders.&nbsp;</p> <p><strong>How to cite:</strong></p> <p>Please refer to the "how to cite" section of the readme file of <a href="https://github.com/astroNils/YOLOv8-BeyondEarth" target="_blank" rel="noopener">https://github.com/astroNils/YOLOv8-BeyondEarth.</a></p> <p><strong>Structure:</strong></p> <pre><code>. └── boulder2024/ ├── jupyter-notebooks/ │ └── REGISTERING_BOULDER_DATASET_IN_DETECTRON2.ipynb ├── test/ │ └── images/ │ ├── &lt;image_name&gt;_image.png │ ├── ... │ └── labels/ │ ├── &lt;image_name&gt;_image.txt │ ├── ... ├── train/ │ └── images/ │ ├── &lt;image_name&gt;_image.png │ ├── ... │ └── labels/ │ ├── &lt;image_name&gt;_image.txt │ ├── ... ├── validation/ │ └── images/ │ ├── &lt;image_name&gt;_image.png │ ├── ... │ └── labels/ │ ├── &lt;image_name&gt;_image.txt │ ├── ... ├── detectron2_inst_seg_boulder_dataset.json ├── README.txt ├── yolo_inst_seg_boulder_dataset.yaml</code></pre> <p>&nbsp;</p> <pre><code>detectron2_inst_seg_boulder_dataset.json</code></pre> <p>is a json file containing the masks as expected by Detectron2 (see <a href="https://detectron2.readthedocs.io/en/latest/tutorials/datasets.html">https://detectron2.readthedocs.io/en/latest/tutorials/datasets.html</a> for more information on the format). In order to use this custom dataset, you need to register the dataset before using it in the training. There is an example how to do that in the jupyter-notebooks folder. You need to have detectron2, and all of its depedencies installed. &nbsp;</p> <pre><code>yolo_inst_seg_boulder_dataset.yaml</code></pre> <p>can be used as it is, however you need to update the paths in the .yaml file, to the test, train and validation folders. More information about the YOLO format can be found here (<a href="https://docs.ultralytics.com/datasets/segment/">https://docs.ultralytics.com/datasets/segment/</a>).</p>

opencc-by-4.0Nov 2024View details →
zenodo52/100

Dominant contribution of Asgard archaea to eukaryogenesis (2024) Tobiasson, V., Koonin, E. PROCESSED DATA AND METADATA

<h1>Main data deposit for "Dominant contribution of Asgard archaea to eukaryogenesis".&nbsp;</h1> <p>Victor Tobiasson, Jacob Luo, Yuri I Wolf, Eugene V Koonin</p> <p>Computational Biology Branch, Division of Intramural Research, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA</p> <p><strong>The Origin of eukaryotes is one of the key problems in evolutionary biology. The demonstration that the Last Eukaryotic Common Ancestor (LECA) already contained the mitochondrion, an endosymbiotic organelle derived from an alphaproteobacterium, and the discovery of Asgard archaea, the closest archaeal relatives of eukaryotes inform and constrain evolutionary scenarios of eukaryogenesis. We undertook a comprehensive analysis of the origins of the core eukaryotic genes tracing to the LECA within a rigorous statistical framework centered around evolutionary hypotheses testing using constrained phylogenetic trees. The results reveal dominant contributions of Asgard archaea to the origin of most of the conserved eukaryotic functional systems and pathways. A limited contribution from Alphaproteobacteria was identified, primarily relating to the energy transformation systems and Fe-S cluster biogenesis, whereas ancestry from other bacterial phyla was scattered across the eukaryotic functional landscape, without consistent trends. These findings suggest a model of eukaryogenesis in which key features of eukaryotic cell organization evolved in the Asgard ancestor, followed by the capture of the Alphaproteobacterial endosymbiont, and augmented by numerous but sporadic horizontal acquisition of genes from other bacteria both before and after endosymbiosis.&nbsp;</strong></p> <div> <div> <div>Version 0.3, updated 180325</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>Main data repository for:</div> <div>Dominant contribution of Asgard archaea to eukaryogenesis (2024)&nbsp;</div> <div>Tobiasson, V., Koonin, E.</div> <div>&nbsp;</div> <div>Contains all final parsed data from the main Eukaryogenesis project&nbsp;</div> <div>investigating the evolutionary ancetries of eukaryotic protein families.&nbsp;</div> <div>&nbsp;</div> <div>Currently (non-static) available at:&nbsp;</div> <div>https://www.biorxiv.org/content/10.1101/2024.10.14.618318v2</div> <div>https://assets-eu.researchsquare.com/files/rs-5352492/v1/2f9c68ae-cf3e-420a-8d29-867b6fb1a878.pdf</div> <div>&nbsp;</div> <div>All code used to generate the data present within this repository available at:&nbsp;</div> <div>https://github.com/VictorTobiasson/eukgen&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>### General information</div> <div>&nbsp;</div> <div>To identify associations between prokaryotic and eukaryotic protein families, separate</div> <div>hidden Markov model (HMM) databases for prokaryotes and eukaryotes were constructed&nbsp;</div> <div>using a custom, cascaded, sequence-to-profile clustering pipeline, implemented using&nbsp;</div> <div>mmseqs2, followed by a multistep data-reduction and multiple sequence alignment (MSA)&nbsp;</div> <div>procedure to generate HMM profiles using hhsuite.&nbsp;</div> <div>&nbsp;</div> <div>A prokaryotic database of 37 million protein sequences was curated from prokaryotic&nbsp;</div> <div>genomes obtained from the NCBI GenBank in November 2023 and supplemented with proteins&nbsp;</div> <div>extracted from 146 Asgard genome assemblies. To avoid inclusion of genes present only&nbsp;</div> <div>within a narrow subset of species, possibly resulting from horizontal transfer from&nbsp;</div> <div>eukaryotes post LECA, we reconstructed the &ldquo;soft-core&rdquo; pangenome for each of the 26&nbsp;</div> <div>curated prokaryotic taxonomic classes. These pangenomes include only those genes that&nbsp;</div> <div>are present in at least 67% of the families within each class of Bacteria and Archaea.&nbsp;</div> <div>The initial eukaryotic database consisted of 30 million protein sequences from 993&nbsp;</div> <div>species taken from EukprotV3 and cleaned using mmseqs2 to remove likely prokaryotic&nbsp;</div> <div>contaminants.&nbsp;</div> <div>&nbsp;</div> <div>Both databases were clustered and MSAs constructed for all non, singleton clusters&nbsp;</div> <div>and HMM profiles created. The resulting eukaryotic HMM dataset was queried against&nbsp;</div> <div>the prokaryotic dataset using hhblits to identify sets of homologous protein sequences.&nbsp;</div> <div>Each eukaryotic cluster and all its significant prokaryotic hits constituted an individual</div> <div>&nbsp;sequence set, hereinafter referred to as an Eukaryotic/Prokaryotic Orthologous Cluster&nbsp;</div> <div>(EPOC). The EPOCs constitute groups of homologous proteins from eukaryotes and prokaryotes&nbsp;</div> <div>(each EPOC contains a unique set of eukaryotic proteins, but some clusters of prokaryotic&nbsp;</div> <div>proteins can be present in multiple EPOCs) that were used for phylogenetic tree&nbsp;</div> <div>construction, annotation, and evolutionary hypothesis testing.&nbsp;</div> <div>&nbsp;</div> <div>To infer the most likely prokaryotic ancestry of the eukaryotic proteins in each EPOC,&nbsp;</div> <div>rather than relying on the tree topology directly, we employed a probabilistic approach&nbsp;</div> <div>for evolutionary hypothesis testing using constraint trees. We exhaustively sampled all&nbsp;</div> <div>arrangements of likely sister clades and obtained Expected Likelihood Weights (ELW) for&nbsp;</div> <div>the set of possible sister clade models. As the ELW metric is analogous to model selection&nbsp;</div> <div>confidence, here we take it to be proportional to the probability of a sampled prokaryotic&nbsp;</div> <div>clade to be the true sister group of the given eukaryotic clade among a set of competing&nbsp;</div> <div>sister clades. For each EPOC, our analysis dynamically accounts for long branch outliers&nbsp;</div> <div>and is robust to phylogenetically non-homogenous clades. This analysis is further capable&nbsp;</div> <div>of resolving eukaryotic paraphyly, treating each eukaryotic clade within a EPOC as a&nbsp;</div> <div>single datapoint for downstream analysis. Our resulting data contains EPOCs annotated&nbsp;</div> <div>using profiles generated from KEGG Orthology Groups (KOGs), each with an MSA generated&nbsp;</div> <div>using muscle5, a maximum likelihood tree inferred using IQtree2 and associated ELW values&nbsp;</div> <div>for all candidate prokaryotic sister phyla. The analysis of prokaryotic ancestry was&nbsp;</div> <div>performed only for those eukaryotic clades that included more than 5 distinct taxonomic&nbsp;</div> <div>labels, with at least one coming from Amorphea and one from Diaphoretickes, the two&nbsp;</div> <div>expansive eukaryotic clades considered to represent either the first or the second&nbsp;</div> <div>bifurcation in the evolution of eukaryotes. Thus, these clades likely represent genes&nbsp;</div> <div>mapping back to the LECA.</div> <div>&nbsp;</div> <div>For further details please see main publication or contact</div> <div>victor.tobiasson@nih.gov</div> <div>eugene.koonin@nih.gov</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>### Included files</div> <div>&nbsp;</div> <div>Unless otherwise stated all files contained are tab separated and utf-8 encoded&nbsp;</div> <div>with the first row containing header information.&nbsp;</div> <div>All data entries encoding lists are &ldquo;|&rdquo; (pipe) separated.&nbsp;</div> <div>Fields without data values are filled with string entries of &ldquo;none&rdquo;.</div> <div>&nbsp;</div> <div>--- Databases ---</div> <div>euk72_ep.tar.gz</div> <div>prok2311_as.tar.gz</div> <div>Prok2311As_final_clusters.tsv</div> <div>Euk72Ep_final_clusters.tsv</div> <div>prok2311_as.hmmDB.tar.gz</div> <div>euk72_ep.hmmDB.tar.gz</div> <div>&nbsp;</div> <div>--- Annotation and Curation ---</div> <div>NCBI_taxonomy_species_addendum.tsv</div> <div>NCBI_taxonomy_class_addendum.tsv</div> <div>Euk72Ep_Prok2311As_final_classes.tsv</div> <div>Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>KEGG_category_mapping.tsv</div> <div>KEGG_metadata.tsv</div> <div>&nbsp;</div> <div>--- EPOC data ---</div> <div>EPOC_data.tar.gz</div> <div>EPOC_annotation_KEGG.tsv</div> <div>EPOC_data.tsv</div> <div>EPOC_data.pangenomes_s10.tsv</div> <div>EPOC_data.pangenomes_s25.tsv</div> <div>EPOC_data.pangenomes_s67.tsv</div> <div>EPOC_data.GTDB.tsv</div> <div>&nbsp;</div> <div># euk72_ep.tar.gz</div> <div>Gunzip-ed .tar archive containing a single directory with 10 files&nbsp;</div> <div>constituting the initial eukaryotic mmseqs2 database with taxonomy annotation.&nbsp;</div> <div>Constructed from a pre-selected list of 72 eukaryotic proteomes downloaded from&nbsp;</div> <div>NCBI as well as a &ldquo;clean&rdquo; version of Eukprot, lacking highly prokaryotic-like&nbsp;</div> <div>contaminant sequences.&nbsp;</div> <div>&nbsp;</div> <div># prok2311_as.tar.gz</div> <div>Gunzip-ed .tar archive containing a single directory with 10 files constituting the&nbsp;</div> <div>initial prokaryotic mmseqs2 database with taxonomy annotation. Constructed from&nbsp;</div> <div>47545 complete genomes retrieved from NCBI in November 2023.&nbsp;</div> <div>&nbsp;</div> <div># prok2311_as.hmmDB.tar.gz</div> <div>Gunzip-ed .tar archive containing 6 files. Comprises an HHSuite Databse formatted&nbsp;</div> <div>from prok2311_as non--singleton clusters, contains 26286 profiles.</div> <div>&nbsp;</div> <div># euk72_ep.hmmDB.tar.gz</div> <div>Gunzip-ed .tar archive containing 6 files. Comprises an HHSuite Databse formatted&nbsp;</div> <div>from euk72_ep non-singleton clusters, contains 1631704 profiles.</div> <div>&nbsp;</div> <div># NCBI_taxonomy_species_addendum.tsv</div> <div>Taxonomy mapping file with manually curated &lsquo;class&rsquo; level annotation for poorly&nbsp;</div> <div>annotated species.&nbsp;</div> <div>&nbsp;</div> <div>taxid: NCBI taxid</div> <div>proposed_class_id: Manually assigned NCBI taxid</div> <div>proposed_class_label: NCBI class name</div> <div>org_name: NCBI organism name</div> <div>&nbsp;</div> <div># NCBI_taxonomy_class_addendum.tsv</div> <div>Class revision file mapping poorly populated class level entries to higher order&nbsp;</div> <div>manually curated labels. Also includes information for small classes with shallow&nbsp;</div> <div>taxonomy which are deleted from the EPOC analysis at the level of tree construction.</div> <div>&nbsp;</div> <div>taxid: NCBI taxid</div> <div>ncbi_class: NCBI taxid of rank corresponding to &lsquo;class&rsquo; following manual&nbsp;</div> <div>amendment as per NCBI_taxonomy_species_addendum.tsv</div> <div>revised_class_id: Manually assigned NCBI taxid of rank corresponding to &lsquo;class&rsquo;</div> <div>revised_class_label: Proposed cleartext name of manually revised revised_class_id&nbsp;</div> <div>&nbsp;</div> <div># Euk72Ep_Prok2311As_final_classes.tsv</div> <div>Final taxonomy at NCBI rank &lsquo;class&rsquo; following revisions for all sequences in Euk72Ep or&nbsp;</div> <div>Prok2311As. These taxonomic labels are used for EPOC tree annotation.&nbsp;</div> <div>&nbsp;</div> <div>acc: mmseqs database header in either prok2311_as or euk72_ep databases</div> <div>taxid: NCBI taxid for organism</div> <div>superkingdom: Top level NCBI taxonomy classification Bacteria, Archaea or Eukarya,&nbsp;</div> <div>used to define Eukaryotic outgroups in EPOC analysis</div> <div>class: Cleartext name of manually revised NCBI rank &lsquo;class&rsquo; identifier for annotation</div> <div>&nbsp;</div> <div># Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>Final taxonomy at GTDB rank &lsquo;phylum&rsquo; transferred using marker genes from GTDB release 220</div> <div>&nbsp;</div> <div>acc: mmseqs database header in either prok2311_as or euk72_ep databases</div> <div>taxid: NCBI taxid for organism</div> <div>superkingdom: Top level NCBI taxonomy classification Bacteria, Archaea or Eukarya,&nbsp;</div> <div>used to define Eukaryotic outgroups in EPOC analysis</div> <div>class: Cleartext name of assigne GTDB phylum</div> <div>&nbsp;</div> <div># Prok2311As_final_clusters.tsv</div> <div>Cluster mapping file for accessions within the initial Prok2311A database to the&nbsp;</div> <div>final clusters used for HMM creation&nbsp;&nbsp;</div> <div>&nbsp;</div> <div>cluster_acc: cluster representative</div> <div>acc: cluster member</div> <div>&nbsp;</div> <div># Euk72Ep_final_clusters.tsv</div> <div>Cluster mapping file for accessions within the initial Prok2311A database to the&nbsp;</div> <div>final clusters used for HMM creation</div> <div>&nbsp;</div> <div>cluster_acc: cluster representative</div> <div>acc: cluster member</div> <div>&nbsp;</div> <div># EPOC_data.tar.gz</div> <div>Gunzip-ed directory containing 16035 EPOC folders. Each folder is named corresponding&nbsp;</div> <div>to the eukaryotic cluster representative which generated its profile as an ID&nbsp;</div> <div>Matches the tree_name field in EPOC_data_prok2311As.tsv</div> <div>contains the following files:</div> <div>&nbsp;</div> <div>&lt;EPOC_ID&gt;.merged.fasta: sequences for all members of the EPOC</div> <div>&lt;EPOC_ID&gt;.merged.fasta.leaf_mapping: tsv separated file containing taxonomy and tree reduction data</div> <div>&lt;EPOC_ID&gt;.merged.fasta.muscle: main cropped MSA for tree generation&nbsp;</div> <div>&lt;EPOC_ID&gt;.merged.fasta.muscle.iqtree: IQtree2 output from tree generation</div> <div>&lt;EPOC_ID&gt;.merged.fasta.muscle.treefile.annot: annotated newick tree file with final tree</div> <div>&lt;EPOC_ID&gt;.merged.tree_data.tsv: final parsed tree data with columns matching&nbsp; EPOC_data_prok2311As.tsv</div> <div>&nbsp;</div> <div>EPOCs with more than one possible eukaryotic sister phyla also contains&nbsp;</div> <div>a folder "constraint_analysis" with constraint tree information used for&nbsp;</div> <div>ELW value calculation.&nbsp;</div> <div>&nbsp;</div> <div># EPOC_data.tsv</div> <div>Main resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs)&nbsp;</div> <div>based on pangenomes defined as including 10% of species per class. This is the main</div> <div>data to be used for genereting the core dataset and for data visualistation</div> <div>Contains information regarding tree breakdown, LCA membership and phylogenetic&nbsp;</div> <div>distances between all detected LCAs. Equivalent to the stacked dataframes from all&nbsp;</div> <div>EPOC directories in EPOC_data&nbsp;</div> <div>&nbsp;</div> <div>tree_name: unique index for each EPOC&nbsp;</div> <div>euk_clade_rep: unique index for each annotated eukaryotic clade within each tree_name</div> <div>euk_clade_size: number of original sequences represented by euk_clade_rep</div> <div>euk_clade_weight: metric for taxonomic purity for each euk_clade_rep</div> <div>euk_leaf_clade: boolean indicating whether euk_clade_rep contains a single leaf</div> <div>euk_LCA: lowest taxa spanning all members in euk_clade_rep</div> <div>euk_scope: list of all taxonomic classes in euk_clade_rep</div> <div>euk_scope_len: length of euk_scope list</div> <div>prok_clade_rep: unique index for each annotated prokaryotic clade for each euk_clade_rep</div> <div>prok_clade_size: number of original sequences represented by prok_clade_rep</div> <div>prok_clade_weight: metric for taxonomic purity for each prok_clade_rep</div> <div>prok_leaf_clade: boolean indicating whether prok_clade_rep contains a single leaf</div> <div>prok_taxa: lowest taxa spanning all members in prok_clade_rep</div> <div>dist: tree-distance from lowest tree node containing all members of prok_clade_rep to lowest tree node containing all members of euk_clade_rep</div> <div>top_dist: graph-distance (node-distance) from lowest tree node containing all members of prok_clade_rep to lowest tree node containing all members of euk_clade_rep</div> <div>raw_stem_length: tree-distance from lowest tree node containing the union of all members of prok_clade_rep and euk_clade_rep to the tree node containing all members of euk_clade_rep</div> <div>median_euk_leaf_dist: median value for all tree distances from the tree node containing all members of euk_clade_rep to the individual leaves</div> <div>stem_length: raw_stem_length/median_euk_leaf_dist</div> <div>logL: log likelihood of best constraint tree constructed</div> <div>deltaL: log likelihood difference between constraint tree for prok_clade_rep and best constraint tree constructed</div> <div>bp-RELL: validation metric from IQtree -trees, see iqtree.org</div> <div>bp-RELL_accept: as above</div> <div>p-KH: as above</div> <div>p-KH_accept: as above</div> <div>p-SH: as above</div> <div>p-SH_accept: as above</div> <div>c-ELW: as above</div> <div>c-ELW_accept: as above</div> <div>p-AU: as above</div> <div>p-AU_accept: as above</div> <div>&nbsp;</div> <div># EPOC_data.pangenomes_s10.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>based on pangenomes defined as including 10% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div>&nbsp;</div> <div># EPOC_data.pangenomes_s25.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>based on pangenomes defined as including 25% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div>&nbsp;</div> <div># EPOC_data.pangenomes_s67.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>based on pangenomes defined as including 67% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div>&nbsp;</div> <div># EPOC_data.GTDB.tsv</div> <div>Resulting data&nbsp; from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>under revised taxonomy from GTDB based on data from Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>Identical file structure to EPOC_data.tsv</div> <div>&nbsp;</div> <div># EPOC_data.alpha_replicates.tsv</div> <div>Resulting data from 20 repetitions of Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>from a subset of Alphaproteobacterial-derived EPOCs.&nbsp;</div> <div>Identical file structure to EPOC_data.tsv with the addition of:</div> <div>&nbsp;</div> <div>rep: indicating technical replicate number, 0-19</div> <div>&nbsp;</div> <div># EPOC_annotation_KEGG.tsv</div> <div>Parsed HHblits output of HMM profiles generated from KEGG KOGs (KEGG Orthologous Groups)&nbsp;</div> <div>against eukaryotic profiles constituting each EPOC</div> <div>&nbsp;</div> <div>Query: query name equal to tree_name from EPOC_data</div> <div>Target: target name equal to kogid in KEGG_category_mapping and KEGG_metadata</div> <div>Prob: data from HHblits, see https://github.com/soedinglab/hh-suite/wiki</div> <div>E-value : as above</div> <div>P-value : as above</div> <div>Score: as above</div> <div>SS: as above</div> <div>Cols: as above</div> <div>Identities: as above</div> <div>Similarity: as above</div> <div>Sum_probs: as above</div> <div>Query-HMM-start: as above</div> <div>Query-HMM-end: as above</div> <div>Template-HMM-start: as above</div> <div>Template-HMM-end: as above</div> <div>Template_columns: as above</div> <div>Template_Neff : as above</div> <div>Pairwise_cov: calculated pairwise coverage from Query and Target start and end</div> <div>Description: category_name from KEGG_category_mapping</div> <div>&nbsp;</div> <div># KEGG_category_mapping.tsv</div> <div>Mapping of relevant KOG identifiers to their higher order categories as&nbsp;</div> <div>"Maps" "Modules" or "Reactions" as per KEGG see https://www.kegg.jp/kegg/pathway.html</div> <div>&nbsp;</div> <div>kogid: unique KOG identifier</div> <div>category_id: KEGG map, module, or reaction number</div> <div>category_name: cleartext name for KOG identifier</div> <div>&nbsp;</div> <div># KEGG_metadata.tsv</div> <div>File mapping KOGs to BRITE classification and to additional databases of chemical properties.</div> <div>&nbsp;</div> <div>kogid: unique KOG identifier</div> <div>name: cleartext name for KOG identifier</div> <div>brite_A: list of BRITE-A sets including KOG</div> <div>brite_B: list of BRITE-A sets including KOG</div> <div>brite_C: list of BRITE-A sets including KOG</div> <div>EC: list of Enzyme commission numbers associated with KOG, see https://enzyme.expasy.org/</div> <div>TC: list of transporter classification numbers associated with KOG, see https://www.tcdb.org/</div> <div>RN: list of KEGG reaction numbers associated with KOG</div> <div>CA: list of CAZY numbers associated with KOG, see http://www.cazy.org/</div> <div>GO: list of GO terms associated with KOG, see https://geneontology.org/</div> </div> <div>&nbsp;</div> </div>

opencc-by-4.0Oct 2024View details →
zenodo52/100

Intermediate processing stage of horizontal particle flux data collected using a snow particle counter on board the R/V Akademik Tryoshnikov in the Southern Ocean during the austral summer of 2016/17 as part of the Antarctic Circumnavigation Expedition (ACE).

<p><strong>Dataset abstract</strong></p> <p>Flux of particles (snow, rain and other particles including sea spray) were recorded passing through a photo-electric snow particle counter installed on board the R/V Akademik Tryoshnikov as part of the Antarctic Circumnavigation Expedition (ACE). Data were recorded from January to March 2017 in the Southern Ocean. Here we present an intermediate step in data processing, with relative horizontal particle flux of particles with a size between 36 &ndash; 2000 &mu;m averaged over one-minute periods. Data are presented in daily files.</p> <p><strong>Dataset contents</strong></p> <ul> <li>SPC_HPF_1min_YYYY_MM_DD.csv, data files, comma-separated values</li> <li>data_file_header.txt, metadata, text</li> <li>README.txt, metadata, text</li> </ul> <p><strong>Dataset license</strong></p> <p>This one-minute averaged horizontal particle flux dataset from ACE is made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0) whose full text can be found at https://creativecommons.org/licenses/by/4.0/</p>

opencc-by-4.0Feb 2021View details →
zenodo52/100

GIXD data of organic-inorganic methylammonium lead bromide perovskite (MAPbBr3), processed q-space maps

<p>This dataset contains grazing incidence x-ray diffraction (GIXD) maps projected in q-space and polar projection. The underlying raw data is published in <a href="https://doi.org/10.5281/zenodo.6683616">10.5281/zenodo.6683616</a> and processed with <a href="https://doi.org/10.5281/zenodo.6683658">10.5281/zenodo.6683658</a>. This data describes a time series of diffraction images acquired with 10 Hz.</p> <p>&nbsp;</p> <p>Parameters of the provided data:</p> <ul> <li> <p>Q-space-maps</p> </li> </ul> <p>&nbsp;</p> <ul> <li> <ul> <li> <p>Horizontal axis (Q<sub>xy</sub>) range: (0, 3.2) &Aring;<sup>-1</sup></p> </li> <li> <p>Vertical axis (Q<sub>z</sub>) range: (0, 3.2) &Aring;<sup>-1</sup></p> </li> <li> <p>Resolution: 1350x1350 pixels</p> </li> <li> <p>Origin (lower left coordinate in q): (0, 0)</p> </li> </ul> </li> <li> <p>Polar data</p> <ul> <li> <p>Horizontal axis (||<strong>q</strong>||) range: (0, 4.53) &Aring;<sup>-1</sup></p> </li> <li> <p>Vertical axis (ф) range: (0, 90) deg</p> </li> <li> <p>Resolution: 512x1024 pixels</p> </li> <li> <p>Origin (lower left coordinate in q): (0, 0)</p> </li> </ul> </li> </ul>

opencc-by-4.0Aug 2022View details →
zenodo52/100

In-situ grazing-incidence X-ray diffraction data of the crystallization process of organic-inorganic methylammonium lead bromide perovskite (MAPbBr3) via employing an isopropanol antisolvent. Raw Data

<p>The dataset contains 400 diffraction images from a 40 second in-situ grazing-incidence wide-angle X-ray scattering measurement of the crystallization process of organic-inorganic methylammonium lead bromide perovskite (MAPbBr3) on a glass substrate. The crystallization is initiated via employing an isopropanol antisolvent during the spin-coating of the perovskite precursor solution. 40 &micro;L of MAPbBr3 solution (4:1 DMF/DMSO solvent mixture) was applied on plasma-cleaned glass substrate in a chamber with kapton windows. The two-phase spin-coating regime included 10 seconds at 1000 rpm followed by 30 seconds at 2000 rpm, 200 &micro;L of antisolvent was dispensed at t = 30 s.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>The data was acquired at the P08 Beamline at PETRA III (DESY Hamburg). Acquisition parameters:</p> <p>&nbsp;</p> <ul> <li> <p>X-ray wavelength: 0.6888 nm</p> </li> <li> <p>Sample detector distance: 809 mm</p> </li> <li> <p>Incidence angle: 0.5 deg.</p> </li> <li> <p>Detector model: XRD 1621 CN3 EHS</p> </li> <li> <p>Acquisition rate&nbsp;: 10 frames per second (10 Hz)</p> </li> <li> <p>Direct beam position (pixels): 545, 222</p> </li> </ul>

opencc-by-4.0Aug 2022View details →
zenodo52/100

Processed glider data: 9 months of hydrographic and ADCP observations in the Gulf of Oman.

<p>68 repeat transects and 2 virtual moorings covering a spring/neap cycle collected by a SeaExplorer glider with T, S, O2, Chl, Optical backscatter, PAR and ADCP data in the Gulf of Oman. Dataset collected as part of the ONR Global project "Shelf slope dyanmics in the Sea of Oman: How submesoscale processes control food and water security". The glider was deployed from the north shore of Oman into the Gulf of Oman, sampling down to 1000m in the oxygen minimum zone.</p> <p>&nbsp;</p> <p>File and variable metadata included in the netCDF files.</p> <p>&nbsp;</p> <p>sea057_M##.ad2cp.#####.nc : Raw ADCP data provided in Nortek .nc format. (version 1.0)</p> <p>SEA057_glider.nc : SeaExplorer data timeseries QC'd and processed into a 1Hz timeseries. (version 1.0)</p> <p>SEA057_ADCP_v2.nc : ADCP data fully processed, binned (2 dbar) and referenced, and then reprojected back onto a timeseries. ADCP data processed as per https://github.com/bastienqueste/gliderad2cp . (version v3)</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo52/100

Liquid Chromatography - Tandem Mass Spectrometry (LC-MS/MS) and Gas Chromatography - Mass Spectrometry (GC-MS) Reference Libraries from Global Natural Products Social Molecular Networking (GNPS) and National Institute of Standards and Technology (NIST) WebBook Processed for Spectral Library Matching

<div>In order to obtain a high-quality LC-MS/MS reference database for spectral library matching, we selected 22 high-quality GNPS tandem mass spectrometry databases generated under the positive ion mode. Further preprocessing similar to Huber et al involving mass-to-charge (m/z) and intensity filtering yields the database found in the file LCMS_GNPS_reference_library.csv which contains 14,705 electrospray ionization (ESI) mass spectra, each of which corresponds to a unique compound. The NIST WebBook database was used to construct GC-MS database contained in the file GCMS_NIST_WebBook.csv. This database contains 23,721 electron ionization (EI) mass spectra, each of which corresponds to a unique non-hyphenated Chemical Abstract Service (CAS) Registry Number.</div> <div>&nbsp;</div> <div>Both LC-MS/MS and GC-MS databases are organized into three columns: one for the identifier, one for the m/z values, and one for the intensity values. For example, if spectrum A has 20 ion fragments, then there will be 20 rows corresponding to spectrum A in the corresponding database with the identifier A repeated 20 times with the corresponding m/z and intensity values.</div>

opencc-by-4.0Jul 2024View details →
zenodo52/100

Labeled Time Series Data of Force/Torque for Monitoring Assembly Processes with a Delta Robot

<p>This dataset comprises 524 recordings of 6-dimensional time series data, capturing forces in three directions and torques in three directions during the assembly of small car model wheels. The data was collected using an equidistant sampling method with a sampling period of 0.004 seconds. Each time series represents the process of assembling one wheel, specifically the placement of a tire onto a rim, and includes a label indicating whether the assembly was successful (OK). The wheels were assembled in batches of four, and the recordings were obtained over six different days. The labels of recordings from two (days 3 and 4) of the six days are invalid as described in [1].&nbsp; The labels presented in this data set are only binary (they do not describe the reason of the failure). The labels of recordings from days 5 and 6 are created by human while the other labels came from a convolutional neural network based computer vision classifier and can be inaccurate as described in section 5.4 of [1].&nbsp; &nbsp;</p> <h4>Dataset Structure:</h4> <ul> <li><strong>File:</strong> <code>ForceTorqueTimeSeries.csv</code> <ul> <li><strong>Columns:</strong> <ul> <li><code>idx (1-524)</code>: Index of the recording corresponding to the assembly of one wheel.</li> <li><code>label (true/false)</code>: Indicates whether the assembly was successful (TRUE = product is OK).</li> <li><code>meas_id (1-6)</code>: Identifier for the day on which the recording was made (refer to Table 2.1 in [1]).</li> <li><code>force_x</code>: X-component of the force measured by the sensor mounted on the delta robot's end effector.</li> <li><code>force_y</code>: Y-component of the force.</li> <li><code>force_z</code>: Z-component of the force.</li> <li><code>torque_x</code>: X-component of the torque.</li> <li><code>torque_y</code>: Y-component of the torque.</li> <li><code>torque_z</code>: Z-component of the torque.</li> </ul> </li> </ul> </li> </ul> <h4>Additional Files:</h4> <ul> <li><strong><code>IMG_3351.MOV</code>:</strong> A video demonstrating the assembly process for one batch of four wheels.</li> <li><strong><code>F3-BP-2024-Trna-Ales-Ales Trna - 2024 - Anomaly detection in robotic assembly process using force and torque sensors.pdf</code>:</strong> Bachelor thesis [1] detailing the dataset and preliminary experiments on fault detection.</li> <li><strong><code>F3-BP-2024-Hanzlik-Vojtech-Anomaly_Detection_Bachelors_Thesis.pdf</code>:</strong> Bachelor thesis [2] describing the data acquisition process.</li> </ul> <h3>References:</h3> <ol> <li>Trna, A. (2024). <em>Anomaly detection in robotic assembly process using force and torque sensors</em> [Bachelor&rsquo;s thesis, Czech Technical University in Prague].</li> <li>Hanzlik, V. (2024). <em>Edge AI integration for anomaly detection in assembly using Delta robot</em> [Bachelor&rsquo;s thesis, Czech Technical University in Prague].</li> </ol>

opencc-by-4.0Sep 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record