Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

558

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

558 results for “Training Data”

Learn how ShareScore rates datasets ↗
zenodo40/100

Training dataset: Generation of a spectral library from HEK-Ecoli Spike-in mass spectrometry data

<p>The five raw files serve as a concise but meaningful training data set in the Galaxy training network (https://galaxyproject.github.io/training-material/).</p> <p>HEK and E.coli cell pellets were lysed with 5 % SDS, 50 mM triethylammonium bicarbonate (TEAB), pH 7.55. The obtained protein extracts were reduced by adding f.c. 5 mM TCEP and alkylated by the addition of f.c. 10 mM iodacetamide. Protein digestion and purification was performed on S-Trap columns. To ensure protein binding to the S-Trap columns, samples were acidified to a final concentration of 1.2 % phosphoric acid (~ pH 2). Six times the sample volume S-Trap buffer (90% aqueous methanol containing a final concentration of 100 mM TEAB, pH 7.1) was added to the samples which were then loaded on the columns and washed with S-Trap buffer. Protein digestion was performed with trypsin and LysC for one hour at 47 &deg;C. Peptides were eluted in three steps with (1) 50 mM TEAB, (2) 0.2 % aqueous formic acid and (3) 50 % acetonitrile containing 0.2 % formic acid. Eluted peptides of HEK and E.coli were mixed in the following ratios (amount in &micro;g):</p> <p>Sample&nbsp;&nbsp; &nbsp;HEK&nbsp;&nbsp; &nbsp;E.coli&nbsp;&nbsp; &nbsp;MS method<br> Sample1&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.00&nbsp; &nbsp; &nbsp; &nbsp; DDA<br> Sample2&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.05&nbsp; &nbsp; &nbsp; &nbsp; DDA<br> Sample3&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.15&nbsp; &nbsp; &nbsp; &nbsp; DDA<br> Sample4&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.40&nbsp; &nbsp; &nbsp; &nbsp; DDA<br> Sample5&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.80&nbsp; &nbsp; &nbsp; &nbsp; DDA</p> <p>Additionally, iRT peptides were added and 1&micro;g of each samples&nbsp;was measured with a Q-Exactive Plus mass spectrometer. Besides the five&nbsp;raw files, we uploaded two&nbsp;fasta files that serve&nbsp;as human and ecoli protein sequence databases, an transition list for the iRT peptides as well as an experimental design for the MaxQuant search.<br> Additionally, we uploaded&nbsp;the Galaxy MaxQuant training result files: protein groups, peptides, mqpar, msms, evidence&nbsp;and PTXQC.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Training dataset: DIA data analysis of a HEK/Ecoli Spike-in dataset using OpenSwathWorkflow

<p>The eight&nbsp;raw files serve as a concise but meaningful training data set in the Galaxy training network (https://galaxyproject.github.io/training-material/).</p> <p>HEK and E.coli cell pellets were lysed with 5 % SDS, 50 mM triethylammonium bicarbonate (TEAB), pH 7.55. The obtained protein extracts were reduced by adding f.c. 5 mM TCEP and alkylated by the addition of f.c. 10 mM iodacetamide. Protein digestion and purification was performed on S-Trap columns. To ensure protein binding to the S-Trap columns, samples were acidified to a final concentration of 1.2 % phosphoric acid (~ pH 2). Six times the sample volume S-Trap buffer (90% aqueous methanol containing a final concentration of 100 mM TEAB, pH 7.1) was added to the samples which were then loaded on the columns and washed with S-Trap buffer. Protein digestion was performed with trypsin and LysC for one hour at 47 &deg;C. Peptides were eluted in three steps with (1) 50 mM TEAB, (2) 0.2 % aqueous formic acid and (3) 50 % acetonitrile containing 0.2 % formic acid. Eluted peptides of HEK and E.coli were mixed in two different ratios and four replicates&nbsp;of each Spike/in ratio were measured:</p> <p>Sample&nbsp;&nbsp; &nbsp;HEK&nbsp;&nbsp; &nbsp;E.coli&nbsp;&nbsp; &nbsp;MS method<br> Sample1&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.15&nbsp; &nbsp; &nbsp; &nbsp; DIA<br> Sample2&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.15&nbsp; &nbsp; &nbsp; &nbsp; DIA<br> Sample3&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.15&nbsp; &nbsp; &nbsp; &nbsp; DIA<br> Sample4&nbsp; &nbsp; 2.5&nbsp; &nbsp; &nbsp; 0.15&nbsp; &nbsp; &nbsp; &nbsp; DIA<br> Sample5&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.80&nbsp; &nbsp; &nbsp; &nbsp; DIA<br> Sample6&nbsp; &nbsp; 2.5&nbsp; &nbsp; &nbsp; 0.80&nbsp; &nbsp; &nbsp; &nbsp; DIA<br> Sample7&nbsp; &nbsp; 2.5&nbsp; &nbsp; &nbsp; 0.80&nbsp; &nbsp; &nbsp; &nbsp; DIA<br> Sample8&nbsp; &nbsp; 2.5&nbsp; &nbsp; &nbsp; 0.80&nbsp; &nbsp; &nbsp; &nbsp; DIA</p> <p>Additionally, iRT peptides were added and 1&micro;g of each samples&nbsp;was measured using&nbsp;data independent acquisition with a Q-Exactive Plus mass spectrometer. Briefly, a scan range from 400-1000 m/Z was first covered by an MS1 scan followed by 25 consecutive MS2 scans (each&nbsp;24 m/z broad). In the next cycle another&nbsp;MS1 scan was acquired followd by 26 MS2 scans (also 24m/z broad) in which the window centers were shifted by 50% compared to the previous cycle of MS2 scans. The resulting raw files contain overlapping MS2 scans.</p> <p>Besides the eight raw files, we uploaded a spectral library, a&nbsp;transition list for the iRT peptides as well as an sample annotation file.<br> Additionally, we uploaded&nbsp;the Galaxy PyProphet score&nbsp;training result files: PyProphet score report and PyProphet score.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Training data for 'Genome annotation with Maker' tutorial (Galaxy Training Material)

<p>The data provided here are part of a Galaxy Training Network tutorial for genome annotation with Maker.</p> <p>It is based on data used in <a href="http://weatherby.genetics.utah.edu/MAKER/wiki/index.php/MAKER_Tutorial_for_WGS_Assembly_and_Annotation_Winter_School_2018">another Maker tutorial</a>.</p> <p>The full genome was <a href="https://www.ncbi.nlm.nih.gov/genome/?term=Schizosaccharomyces%20pombe[Organism]&amp;cmd=DetailsSearch">downloaded from NCBI</a>, and mitochondria sequence removed from it for simplicity.</p>

opencc-by-4.0Aug 2018View details →
zenodo40/100

Doctoral Studies as part of an Innovative Training Network (ITN): Early Stage Researcher (ESR) experiences - supplemental material & data

<p>Table and data repository for the manuscript &quot;Doctoral Studies as part of an Innovative Training Network (ITN): Early Stage Researcher (ESR) experiences&quot;</p> <p><strong>Supplemental Tables:</strong></p> <ul> <li>table1_ESIT Project Table</li> <li>table2_TIN-ACT Project Table</li> <li>table3_ITN_tinnitus</li> <li>table4_ITN_other</li> <li>table5_Individual PhDs</li> </ul> <p><strong>Individual-level and de-identified survey data (raw data):</strong></p> <ul> <li>raw_data_ITN_tinnitus (survey results from PhDs as part of an ITN with a focus on tinnitus)</li> <li>raw_data_ITN_other (survey results from PhDs associated to ITNs with another focus)</li> <li>raw_data_Individual_Phds (survey results from PhDs not part of an ITN)</li> </ul>

opencc-by-4.0Jan 2021View details →
zenodo40/100

H2M survey data on commercialisation training needs of Health Researchers

<p>Health-2-Market was a 3-year long Coordination and Support Action, funded by the European Union&rsquo;s Seventh Framework Programme for research, technological development and demonstration (Grant Agreement No 305532). H2M aimed at providing training and individual support to Health / Life Sciences researchers in the process of translating their research results into successful new business ideas.</p> <p>With a view to properly adapting the training offer of the project to the needs of Health / Life Sciences researchers in terms of entrepreneurship and business skill development a Training Needs Analysis (TNA) was conducted. In this context, H2M launched an online survey targeted at Health / Life Sciences researchers who have been involved in EU health projects. In particular, the objectives of the survey were:</p> <ul> <li>To formulate&nbsp; a&nbsp; descriptive&nbsp; understanding&nbsp; of various&nbsp; aspects&nbsp; of&nbsp; commercialisation&nbsp; and&nbsp; training&nbsp; needs&nbsp; of &nbsp;the main target group of the project;</li> <li>To divide this target group into homogeneous sub-groups (clusters) along a number of key characteristics such as demographics, commercialisation attitudes and needs;</li> <li>To understand preferences and importance of different aspects and needs through the analysis of: <ul> <li>Knowledge&nbsp; areas&nbsp; that&nbsp; can&nbsp; influence&nbsp; commercialisation behaviour;</li> <li>Training modalities&nbsp; that&nbsp; have&nbsp; an&nbsp; effect&nbsp; on&nbsp; the&nbsp; intention&nbsp; to participate and /or on the perception of the usefulness of a commercialisation training;</li> <li>Variations identified over different sub-groups.</li> </ul> </li> </ul> <p>The survey was dispatched to a database composed of 7,991 unique contacts of participants in previous health projects, accessed through the Directorate General for Health and Food Safety of the European Commission. The initial aim of at least 50 complete responses was overwhelmingly surpassed: 637 respondents completed the survey in full.</p> <p>The &ldquo;H2M survey data on commercialisation training needs of Health Researchers&rdquo; dataset contains the raw, anonymised data that were collected from these respondents, along with the questionnaire items that were utilised.</p>

opencc-by-nc-4.0Sep 2015View details →
zenodo40/100

Survey data for "Remote Sensing & GIS Training in Ecology and Conservation"

<p>This file provides the raw data of an online survey intended at gathering information regarding remote sensing (RS) and Geographical Information Systems (GIS) for conservation in academic education. The aim was to unfold best practices as well as gaps in teaching methods of remote sensing/GIS, and to help inform how these may be adapted and improved. A total of 73 people answered the survey, which was distributed through closed mailing lists of universities and conservation groups.</p>

opencc-zeroApr 2016View details →
zenodo40/100

A rule based Tibetan part-of-speech (POS) tagger for the creation of gold standard training data

<p>This rule based Tibetan part-of-speech (POS) tagger was prepared in the course of the research project &#39;Tibetan in Digital Communication&#39; (2012-2015) hosted at SOAS, University of London and funded by the UK&#39;s Arts and Humanities Research Council (grant code: AH/J00152X/1). For a description of the tag set see Garrett et al. 2014. and Garrett et al. 2015. For a description of the tagger itself see Garrett et al. 2014. Note that the tagger must be used together with a lexicon (for example Hill &amp; Garrett 2017a). One must use one&#39;s own script to tag all words with all tags in the lexicon and then apply the tagger to remove incorrect tags.</p> <p>On the associated corpus of 318,230 words (Hill &amp; Garrett 2017b) the lexical tagger (i.e. simply applying all available tags to all words) tags 141,911 words with the correct unique tag, achieves as accuracy of 1.000 (by definition getting the right tag among others for each word) with an ambiguity of 2.73111. In contrast, the Rule Tagger tags 241,256 words with the correct unique tag, achieves an accuracy of 0.99893 and an ambiguity of 1.38577.</p> <p>Because this tagger does not achieve ambiguity 1.000 it is not suitable for tagging large scale corpora, but instead is useful for the creation of gold standard training data.</p> <p>N.B. In some rare cases the tagger removes all POS-tags.</p>

opencc-by-4.0May 2017View details →
zenodo40/100

Training material for small RNA-seq data analysis (Galaxy Training Network tutorial)

<p>The data provided here are part of a Galaxy Training Network tutorial that analyzes small RNA-seq (sRNA-seq) data from a study published by Harrington et al. (DOI:10.1186/s12864-017-3692-8) to detect differential abundance of various classes of endogenous short interfering RNAs (esiRNAs). The goal of this study was to investigate "connections between differential retroTn and hp-derived esiRNA processing and cellular location, and to investigate the potential link between mRNA 3’ end cleavage and esiRNA biogenesis." To this end, sRNA-seq libraries were constructed from triplicate <em>Drosophila</em> tissue culture samples under conditions of either control RNAi or RNAi knockdown of a factor involved in mRNA 3’ end processing, <em>Symplekin</em>. This dataset (GEO Accession: GSE82128) consists of single-end, size-selected, non-rRNA-depleted sRNA-seq libraries. Because of the long processing time for the large original files, we have downsampled the original raw data files to include only reads that align to a subset of interesting transcript features including: (1) transposable elements, (2) <em>Drosophila</em> piRNA clusters, (3) <em>Symplekin</em>, and (4) genes encoding mass spectrometry-defined protein binding partners of <em>Symplekin</em> from Additional File 2 in the indicated paper by Harrington et al. More details on features 1 and 2 can be found here: https://github.com/bowhan/piPipes/blob/master/common/dm3/genomic_features (piRNA_Cluster, Trn). All features are from the <em>Drosophila</em> genome Apr. 2006 (BDGP R5/<em>dm3</em>) release.</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

Training data for MaxQuant and Msstats label-free analysis in Galaxy

<p>The files serve as input and intermediate results for a MaxQuant and Msstats training on skin cancer tissues (<a href="https://doi.org/10.1016/j.matbio.2017.11.004">https://doi.org/10.1016/j.matbio.2017.11.004</a>) in the Galaxy training network (https://training.galaxyproject.org).</p> <p>Input files: human FASTA database for Maxquant. Annotation file and comparison matrix file for Msstats.</p> <p>Intermediate result files: MaxQuant protein groups, evidence and PTXQC.</p>

opencc-by-4.0Feb 2021View details →
zenodo40/100

Training and test data, plus saved models for the upcoming paper `Top-down perceptual inference shaping the activity of early visual cortex'

<p>Each .pkl&nbsp;file contains a training or test dataset&nbsp;in the form of a Python dictionary (generated with Python 3.8.5) with the following fields:</p><ul><li>'train_images': 640,000 float32 images&nbsp;used&nbsp;for model training. These are 40px images that contain 1600 pixel intensities each.</li><li>'train_labels': float32 labels for each image in&nbsp;'train_images'. All natural images are&nbsp;labeled&nbsp;with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0,&nbsp;according to their texture family.</li><li>'test_images': 64,000 float32 images&nbsp;used&nbsp;for model testing.&nbsp;These are 40px images that contain 1600 pixel intensities each.</li><li>'test_labels': float32 labels for each image in&nbsp;'test_images'. All natural images are&nbsp;labeled&nbsp;with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0,&nbsp;according to their texture family.</li></ul><p>The .zip file contains a saved model snapshot and various intermediate evaluative data.&nbsp;Details on these are coming soon.</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Lund Bladder Cancer Group - Lund Taxonomy 2023 Classifier - Training Data

<p>Raw and processed training data used for the Lund Taxonomy 2023 gene expression classifier for urothelial carcinoma.</p><p>The dataset contains the raw .CEL files for 3 Affymetrix Gene 1.0 ST cohorts: Lund2017 (n=307, GSE83586), Lund2020 (n=173, GSE128959), Lund2022 (n=310, 117 from GSE169455, 37 from GSE222073, and 156 previously unpublished samples). Code examples for RMA and SCAN.UPC normalization for Affymetrix Gene, Exon, and HTA2 arrays used in the study is included. Hybridization batch information is available for each array cohort.</p><p>The dataset contains Kallisto and Salmon transcript quantification files for 265 RNA-sequenced tumors and the tx2gene file+code for transcript to gene summarization using txImport</p><p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Improving Algorithm-Selectors and Performance-Predictors via Learning Discriminating Training Samples - Code and Data

<p>This repository contains the code and data for reproducibility of the paper 'Improving Algorithm-Selectors and Performance-Predictors via Learning Discriminating Training Samples'.&nbsp;</p> <p>The following files are included:</p> <ul> <li>Plots: Additional plots not in the paper;</li> <li>Code: Python scripts to generate trajectories and perform classification/regression;</li> <li>best_algo.csv : Labels for the classification;</li> <li>performances.csv : Performances used for the regression;</li> <li>SA_parameters.csv : SA parameters for all machine learning tasks;</li> <li>irace_scenario.txt : scenario used for the tuning.</li> </ul>

opencc-by-4.0Jan 2024View details →
dryad40/100

Data and trained models for: Human-robot facial co-expression

<p>Large language models are enabling rapid progress in robotic verbal communication, but nonverbal communication is not keeping pace. Physical humanoid robots struggle to express and communicate using facial movement, relying primarily on voice. The challenge is twofold: First, the actuation of an expressively versatile robotic face is mechanically challenging. A second challenge is knowing what expression to generate so that they appear natural, timely, and genuine. Here we propose that both barriers can be alleviated by training a robot to anticipate future facial expressions and execute them simultaneously with a human. Whereas delayed facial mimicry looks disingenuous, facial co-expression feels more genuine since it requires correctly inferring the human's emotional state for timely execution. We find that a robot can learn to predict a forthcoming smile about 839 milliseconds before the human smiles, and using a learned inverse kinematic facial self-model, co-express the smile simultaneously with the human. We demonstrate this ability using a robot face comprising 26 degrees of freedom. We believe that the ability co-express simultaneous facial expressions could improve human-robot interaction.</p>

opencc-zeroMar 2024View details →
zenodo40/100

AI4Life-MDC24 Challenge data: Fluorescence Microscopy Datasets for Training Deep Neural Networks

<p>This is a subset of the Supporting data for <em>Guy M Hagen, Justin Bendesky, Rosa Machado, Tram-Anh Nguyen, Tanmay Kumar, Jonathan Ventura, Fluorescence microscopy datasets for training deep neural networks, GigaScience, Volume 10, Issue 5, May 2021, giab032, <a href="https://doi.org/10.1093/gigascience/giab032">https://doi.org/10.1093/gigascience/giab032</a></em><br><br>The selected <strong>subset</strong> contains 79 images from Data Set 4 in the form of a single tiff file.&nbsp;<br><br>The paper describing the original dataset is available here: <a href="https://academic.oup.com/gigascience/article/10/5/giab032/6269106">https://academic.oup.com/gigascience/article/10/5/giab032/6269106</a><br>The original dataset is available here: <a href="http://gigadb.org/dataset/100888">http://gigadb.org/dataset/100888</a></p> <p><br>AI4Life has received funding from the European Union&rsquo;s Horizon Europe research and innovation programme under grant agreement number 101057970. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Council Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

IGN Train and Validation Data for ICDAR'24 MapText Competition

<p>Data set of 2Kx2K image tiles cropped from Napoleonic Cadastre maps of the&nbsp;<a href="https://archives.valdemarne.fr/recherches/archives-en-ligne/cadastre-napoleonien">Val de Marne Archive</a> for the <a href="https://rrc.cvc.uab.es/?ch=28">ICDAR'24 Competition on Historical Map Text Detection, Recognition, and Linking</a>.</p> <p>Annotations and images follow the format described at the competition website and can be evaluated using the official <a href="https://github.com/icdar-maptext/evaluation">evaluation repository</a> script.</p> <table> <tbody> <tr> <td>&nbsp;</td> <td><strong>Train</strong></td> <td><strong>Validation</strong></td> </tr> <tr> <td>Annotations</td> <td><code>ign_train.json</code></td> <td><code>ign_val.json</code></td> </tr> <tr> <td>Images</td> <td><code>train.zip</code></td> <td><code>val.zip</code></td> </tr> <tr> <td>Files</td> <td><code>ign/train/*.jpg</code></td> <td><code>ign/val/*.jpg</code></td> </tr> <tr> <td>Tiles</td> <td>80</td> <td>15</td> </tr> <tr> <td>Map Sheets</td> <td>37</td> <td>9</td> </tr> <tr> <td>Words</td> <td>8,096</td> <td>1,801</td> </tr> <tr> <td>Label Groups</td> <td>7,449</td> <td>1,661</td> </tr> <tr> <td>Illegible Words</td> <td>563</td> <td>217</td> </tr> <tr> <td>Truncated Words</td> <td>371</td> <td>91</td> </tr> <tr> <td>Valid Words</td> <td>7,533</td> <td>1,584</td> </tr> </tbody> </table> <p><em>Original images available at <a href="https://archives.valdemarne.fr/recherches/archives-en-ligne/cadastre-napoleonien" target="_blank" rel="noopener">https://archives.valdemarne.fr/recherches/archives-en-ligne/cadastre-napoleonien</a> as of 1 Feb. 2024.</em></p>

opencc-by-sa-4.0Feb 2024View details →
zenodo40/100

Silva taxonomic training data formatted for DADA2 (Silva version 138.2)

<p>These DADA2-formatted training fasta files were derived from the Silva Project's version 138.2 release: https://www.arb-silva.de/</p> <p>These fastas were generated by the following commands (using the dada2 R package version 1.35.4):</p> <blockquote> <p>path &lt;- "~/tax/Silva/v138_2"<br>fn.out.slv &lt;- "~/Desktop/silva_nr99_v138.2_toGenus_trainset.fa.gz"<br>dada2:::makeTaxonomyFasta_SilvaNR(file.path(path, "SILVA_138.2_SSURef_NR99_tax_silva.fasta.gz"),&nbsp;<br>&nbsp; &nbsp; file.path(path, "tax_slv_ssu_138.2.txt"),&nbsp;<br>&nbsp; &nbsp; fn.out.slv)<br>&nbsp; &nbsp;<br>fn.out.spc.slv &lt;- "~/Desktop/silva_nr99_v138.2_toSpecies_trainset.fa.gz"<br>dada2:::makeTaxonomyFasta_SilvaNR(file.path(path, "SILVA_138.2_SSURef_NR99_tax_silva.fasta.gz"),&nbsp;<br>&nbsp; &nbsp; file.path(path, "tax_slv_ssu_138.2.txt"),&nbsp;<br>&nbsp; &nbsp; fn.out.spc.slv, include.species=TRUE)</p> <p>fn.out.aS.slv &lt;- "~/Desktop/silva_v138.2_assignSpecies.fa.gz"<br>dada2:::makeSpeciesFasta_Silva("~/tax/silva/v138_2/SILVA_138.2_SSURef_tax_silva.fasta.gz",&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;fn.out.aS.slv)</p> </blockquote>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Greengenes2 training data formatted for DADA2 (Greengenes2 release version 2024.09)

<p>These DADA2-formatted training fasta files were derived from the Greengenes2 version 2024.09 release. https://ftp.microbio.me/greengenes_release/2024.09/</p> <p>These fastas were generated by the following commands (using the dada2 R package version 1.35.4):</p> <blockquote> <p>path &lt;- "~/tax/GG2/2024_09"<br>fn &lt;- file.path(path, "5b42d9b6-2f24-4f01-b989-9b4dafca7d5e/data/dna-sequences.fasta")<br>txfn &lt;- file.path(path, "b7c3e691-ea51-4547-94dd-f79f49e41a36/data/taxonomy.tsv")</p> <p>fn.out.gg &lt;- "~/Desktop/gg2_2024_09_toGenus_trainset.fa.gz"<br>dada2:::makeTaxonomyFasta_GG2(fn, txfn, fn.out.gg, include.species=FALSE, compress=TRUE)</p> <p>fn.out.spc.gg &lt;- "~/Desktop/gg2_2024_09_toSpecies_trainset.fa.gz"<br>dada2:::makeTaxonomyFasta_GG2(fn, txfn, fn.out.spc.gg, include.species=TRUE, compress=TRUE)</p> </blockquote>

opencc-by-4.0Nov 2024View details →
zenodo40/100

RDP taxonomic training data formatted for DADA2 (RDP release 19 - update 2023-08-23)

<p>These DADA2-formatted training fasta files were derived from the Ribosomal Database Project's Training Set 19 and the 2023-08-23 release of the RDP database. https://sourceforge.net/projects/rdp-classifier/files/RDP_Classifier_TrainingData/</p> <p>These fastas were generated by the following commands using the dada2 R package version 1.35.4:</p> <blockquote> <p>## RDP data: https://sourceforge.net/projects/rdp-classifier/files/RDP_Classifier_TrainingData/</p> <p>path &lt;- "~/tax/rdp/v19"<br>fn.out.rdp &lt;- "~/Desktop/rdp_19_toGenus_trainset.fa.gz"<br>dada2:::makeTaxonomyFasta_RDP(file.path(path, "trainset19_072023_speciesrank.fa"),&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;file.path(path, "trainset19_db_taxid.txt"),&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;fn.out.rdp, include.species=FALSE,<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;compress=TRUE)</p> <p>fn.out.spc.rdp &lt;- "~/Desktop/rdp_19_toSpecies_trainset.fa.gz"<br>dada2:::makeTaxonomyFasta_RDP(file.path(path, "trainset19_072023_speciesrank.fa"),&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;file.path(path, "trainset19_db_taxid.txt"),&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;fn.out.spc.rdp, include.species=TRUE,<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;compress=TRUE)</p> </blockquote>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Galaxy Training Data for "Overview of the Galaxy OMERO-suite"

<div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <p><strong>Sub-sample images:&nbsp;</strong> Images of cytoplasm to nucleus translocation of the transcription factor NF&kappa;B in MCF7 (human breast adenocarcinoma cell line) and A549 (human alveolar basal epithelial) cells in response to TNF&alpha; concentration. Images are at 10x objective magnification. The plate was acquired at Vitra Bioscience on the CellCard reader. For each well there is one field with two images: a nuclear counterstain (DAPI) image and a signal stain (FITC) image. Image size is 1360 x 1024 pixels. Images are in 8-bit BMP format.</p> <p><strong>Source</strong>: https://bbbc.broadinstitute.org/BBBC014</p> <p><strong>Citation</strong>: Ljosa, Vebjorn, Katherine L. Sokolnicki, and Anne E. Carpenter. "Annotated high-throughput microscopy image sets for validation." Nature methods 9.7 (2012): 637-637. <span>https://doi.org/10.1038/nmeth.2083</span></p> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Processed data and trained models for "HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields"

<p>#############</p> <p>HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields, CVPR 2024</p> <p>#############</p> <p>Haozhe Qi, Chen Zhao, Mathieu Salzmann, Alexander Mathis.</p> <p>Affiliation: EPFL</p> <p>Date: June, 2024</p> <p>Link to the CVPR article: <a href="https://openaccess.thecvf.com/content/CVPR2024/papers/Qi_HOISDF_Constraining_3D_Hand-Object_Pose_Estimation_with_Global_Signed_Distance_CVPR_2024_paper.pdf">https://openaccess.thecvf.com/content/CVPR2024/papers/Qi_HOISDF_Constraining_3D_Hand-Object_Pose_Estimation_with_Global_Signed_Distance_CVPR_2024_paper.pdf</a></p> <p>Link to the Arxiv article: <a href="https://arxiv.org/abs/2402.17062">https://arxiv.org/abs/2402.17062</a></p> <p>--------------------------------</p> <div> <div>Here we provide the data of our article "HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields". It contains the preprocessed data of the interacting objects and SDF samples. Meanwhile, we also include the trained model weights here.</div> <br> <div>The overall structure of the data is:</div> <br> <div>├── <a href="../api/records/11668766/draft/files/ckpts.zip/content" target="_blank" rel="noopener noreferrer">ckpts.zip</a>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - Contains the trained weights model on different datasets (DexYCB and HO3Dv2)</div> <div>├── <a href="../api/records/11668766/draft/files/annotations.zip/content" target="_blank" rel="noopener noreferrer">annotations.zip</a>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;- Contains the preprocessed annotations of DexYCB and HO3Dv2 for efficient data loading.</div> <div>├── <a href="../api/records/11668766/draft/files/simple_ycb_models.zip/content" target="_blank" rel="noopener noreferrer">simple_ycb_models.zip</a>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;- Contains the preprocessed YCB objects for batched evaluation.</div> <div>├── <a href="../api/records/11668766/draft/files/test.zip/content" target="_blank" rel="noopener noreferrer">test.zip</a>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - Contains the processed SDF files for DexYCB test set.</div> <div>├── <a href="https://zenodo.org/api/records/14190951/draft/files/ho3d_release.zip/content" target="_blank" rel="noopener noreferrer">ho3d_release.zip</a>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;- Contains the HO3Dv2 submission trained with HO3D training set.</div> <div>├── <a href="https://zenodo.org/api/records/14190951/draft/files/ho3d_render_release.zip/content" target="_blank" rel="noopener noreferrer">ho3d_render_release.zip</a>&nbsp; &nbsp; &nbsp; &nbsp;- Contains the HO3Dv2 submission trained with HO3D training set and rendering set.</div> <div>&nbsp;</div> <br> <div>The code to reproduce the results is available at: <a href="https://github.com/amathislab/HOISDF">https://github.com/amathislab/HOISDF</a></div> <div>&nbsp;</div> <div>--------------------------------</div> </div> <p>If you find our code, weights, predictions or ideas useful, please cite:</p> <p>@inproceedings{qi2024hoisdf,<br>&nbsp; title={HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields},<br>&nbsp; author={Qi, Haozhe and Zhao, Chen and Salzmann, Mathieu and Mathis, Alexander},<br>&nbsp; booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},<br>&nbsp; pages={10392--10402},<br>&nbsp; year={2024}<br>}</p>

opencc-by-4.0Jun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record