Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
392
datasets available to search
ShareScore release 0.7.1
Dataset results
392 results for “Tutorial”
Pylake tutorial dataset: Population dynamics
<p>A dataset of force data measured on a LUMICKS C-Trap. The system consists of a DNA hairpin construct tethered between two polystyrene beads held at a constant distance with the optical traps.</p>
Pylake tutorial dataset: Force extension curves
<p>Force extension curves of various constructs acquired on a LUMICKS C-Trap.</p> <p>The constructs are</p> <ul> <li>a DNA Hairpin</li> <li>Adenylate Kinase with DNA handles</li> <li>DNA in the presence and absence of RecA</li> </ul>
Pylake tutorial dataset: Scans
<p>2D confocal scans of protein binding acquired on a LUMICKS C-Trap. The system consists of </p> <ul> <li>lambda DNA with two Atto647N fluorophores attached at specific locations (red channel)</li> <li>protein labeled with an Atto565 fluorophore (green channel)</li> </ul>
Cryptosporidium IFA tutorial
<p>Video tutorial for detection of<em> Cryptosporidium</em> spp. oocysts in fecal specimens.</p>
Cryptosporidium Nested-PCR on 18S rDNA Tutorial
<p>Video tutorial for<em> Cryptosporidium</em> spp. identification by nested-PCR on 18S rDNA locus.</p>
Data Set For Revels-MD tutorials
<p>This small data set contains the trajectory files necessary to run the tutorials for the revelsmd (<a href="https://github.com/user200000/revelsmd">https://github.com/user200000/revelsmd</a>) the trajectories were generated using lammps (<a href="https://www.lammps.org/#gsc.tab=0">https://www.lammps.org/#gsc.tab=0</a>) and gromacs (<a href="https://www.gromacs.org/">https://www.gromacs.org/</a>). </p> <ul> <li>Lennard jones sphere radial distribution functions (as in <a href="https://doi.org/10.1063/5.0053737">https://doi.org/10.1063/5.0053737</a>) (number 1)</li> <li>Solvation of an immobilised Lennard jones sphere in a solvent of identicle Lennard Jones spheres.(number 2)</li> <li>Solvation of a static water molecule (as in <a href="https://aip.scitation.org/doi/abs/10.1063/1.5111697">https://aip.scitation.org/doi/abs/10.1063/1.5111697</a>) (number 4)</li> </ul> <p><br> A fourth tutorial is in development</p>
openeddy_tutorials
<p>The example dataset and R Markdown file needed to run the openeddy R package tutorials. The updated version of R Markdown file with tutorials can be found at <a href="https://github.com/lsigut/openeddy_tutorials">https://github.com/lsigut/openeddy_tutorials</a>. The example data set is stored here to overcome the file size limitation of GitHub.</p>
Training Data for "Binning of metagenomic sequencing data" tutorial
<p><strong>Metagenomics is the study of genetic material recovered directly from environmental samples, such as soil, water, or gut contents, without the need for isolation or cultivation of individual organisms. Metagenomics binning is a process used to classify DNA sequences obtained from metagenomic sequencing into discrete groups, or bins, based on their similarity to each other</strong>. The goal of metagenomics binning is to assign the DNA sequences to the organisms or taxonomic groups that they originate from, allowing for a better understanding of the diversity and functions of the microbial communities present in the sample. This is typically achieved through computational methods that use sequence similarity, composition, and other features to group the sequences into bins.</p> <p>There are two main types of metagenomics binning: <strong>reference-based</strong> and <strong>de novo</strong>.</p> <ul> <li><strong>reference-based binning</strong> involves aligning the sequences to a database of known genomes or reference sequences</li> <li><strong>de novo binning</strong> involves clustering the sequences based on similarity without prior knowledge of the organisms or reference sequences present in the sample.</li> </ul> <p>Both methods have their strengths and limitations, and researchers often use a combination of approaches to improve the accuracy of their binning results. Metagenomics binning is an important tool for understanding the functional potential of microbial communities in various environments and has applications in fields such as biotechnology, environmental science, and human health.</p> <p>In this tutorial, we will learn how to run metagenomic binning tools and evaluate the quality of the results. In order to do that, we will use data from the study: <a href="https://www.ebi.ac.uk/metagenomics/studies/MGYS00005630#overview">Temporal shotgun metagenomic dissection of the coffee fermentation ecosystem</a> and MetaBAT2 algorithm. For an in-depth analysis of the structure and functions of the coffee microbiome, a temporal shotgun metagenomic study (six time points) was performed. The six samples have been sequenced with Illumina MiSeq utilizing whole genome sequencing.</p> <p>Based on the 6 original dataset of the coffee fermentation system, we generated mock datasets for this tutorial.</p>
Training data for 'Genome annotation with Funannotate' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for genome annotation with funannotate.</p> <p>Genome was assembled following the GTN Flye assembly tutorial, then masked with RepeatMasker.</p> <p>RNASeq data: SRR8534859 reads were mapped to the genome using STAR (toolshed.g2.bx.psu.edu/repos/iuc/rgrnastar/rna_star/2.7.8a+galaxy0), then the bam was downsampled (10% with toolshed.g2.bx.psu.edu/repos/devteam/picard/picard_DownsampleSam/2.18.2.1) to reduce the size of the dataset. Fastq files were then extracted from the resulting bam file (toolshed.g2.bx.psu.edu/repos/devteam/picard/picard_SamToFastq/2.18.2.1).</p> <p>SwissProt_subset.fasta is a subset of SwissProt proteins that are known to have some similarity with the genome (found using Diamond against the genome, then extracting sequences matching with e-value < 0.0001).</p>
Video Tutorial for the FAIR Game
<p>A video tutorial about the FAIR Game</p>
Dataset to be used with SEPIA tutorial
<p>This data collection contains a multi-echo dataset to be used in tutorials from the <a href="https://github.com/kschan0214/sepia">SEPIA software</a>. The data corresponds to a modification of one of the datasets distributed with the QSM harmmonization effort that can be found in - https://zenodo.org/record/7795492.</p> <p>The tutorial associated with this dataset can be found on the <a href="https://sepia-documentation.readthedocs.io/en/latest/">documentation website</a> and is referred to as Toolkit 2023.</p> <p> </p> <p> </p>
Immcantation 10x Tutorial Data
<p>Necessary datasets to run the Immcantation 10x Tutorial. Below is the description of the files in the data set. </p> <ul> <li>BCR_data_sample1.tsv: data corresponding to the first sample (sample 1) of the two samples analyzed in the 10x tutorial. This is the sample used to show the Change-O steps.</li> <li>filtered_contig_annotations.csv: filtered contig annotations file for sample 1, output of cellranger vdj.</li> <li>filtered_contig.fasta: sequence fasta file for sample 1, output of cellranger vdj.</li> <li>BCR_data.tsv: AIRR rearrangement file containing the data for both samples 1 and 2 used in the 10x tutorial.</li> <li>BCR.data_08112023.rds: R dataframe object containing the single-cell BCR sequencing data for both samples 1 and 2 used in the 10x tutorial.</li> <li>GEX.data_08112023.rds: Seurat object containing the single-cell gene expression data used in the 10x tutorial.</li> </ul>
Source code for R tutorials and dataset for empirical case study on Malurus elegans (red-winged fairy wren)
Open the record for dataset details and reuse information.
Tutorial video for: A toolbox for the retrodeformation and muscle reconstruction of fossil specimens in Blender
Open the record for dataset details and reuse information.
Supplementary datasets, data analysis code, and R tutorials for: Phylogenetic analysis of adaptation in comparative physiology and biomechanics: overview and a case study of thermal physiology in treefrogs
Open the record for dataset details and reuse information.
Training data for 'Unicycler assembly of SARS-CoV-2 genome with preprocessing to remove human genome reads' tutorial (Galaxy Training Material)
<p>The data here is a copy of the corresponding SRR records in the NCBI SRA. The duplication serves a dual purpose:</p> <ol> <li>as a backup should there be problems connecting to NCBI servers, e.g., during Galaxy user trainings.</li> <li>to illustrate how to obtain raw sequencing data from alternative sources, and to organize the data into the same collection structure in a Galaxy history that is generated by specialized Galaxy SRA download tools.</li> </ol>
Datasets for GTN tutorial on SARS-CoV-2 variant analysis
<p>A reference genome in FASTA format is provided for SARS-CoV-2, "Severe acute respiratory syndrome coronavirus 2 isolate Wuhan-Hu-1, complete genome", having the accession ID of NC_045512.2.</p> <p>This file was obtained from NCBI within this Galaxy history: https://usegalaxy.org/u/dan/h/nc0455122-from-ncbi</p>
IPBES Data Management Tutorials - Session 4.2: Recommendations for workflow establishment and data management
<p>The <em>IPBES data management tutorials</em> are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The <em>Data management of active research data </em>chapter provides an introduction for IPBES experts on how to manage data while actively being used, analyzed, and produced to fulfill the criteria of the IPBES data management policy.</p> <p>This session,<em> Recommendations for workflow establishment and data management</em>, provides recommendations for data management including the generation of workflows and data storage. </p>
IPBES Data Management Tutorials - Session 4.1: General introduction to the management of active research data
<p>The <em>IPBES data management tutorials</em> are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The <em>Data management of active research data </em>chapter will provide an introduction for IPBES experts on how to manage data while actively being used, analyzed, and produced to fulfill the criteria of the IPBES data management policy.</p> <p>This session,<em> General introduction to the management of active research data</em>, covers what active data management is, why it is important, and considerations for key decisions made in this process. </p>
IPBES Data Management Tutorials - Session 3.1: General introduction to data management reports
<p>The <em>IPBES data management tutorials</em> are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The <em>IPBES data management reports </em>chapter provides an overview and discussion of specific elements of IPBES data management reports.</p> <p>This session provides a <em>general introduction to data management reports</em> and explains the difference between data management plans and data management reports.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.