Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,582
datasets available to search
ShareScore release 0.9.0
Dataset results
1,582 results for “manuscript”
Supplementary Data for Manuscript 'Observing impacts on luminescence depth profile evolutions from surface altered quartzite using OSL laser scanning and controlled light exposed rock sampling techniques'
<p>This file contains the supplementary data for the manuscript 'Observing impacts on luminescence depth profile evolutions from surface altered quartzite using OSL laser scanning and controlled light exposed rock sampling techniques'</p>
Representative data accompanying the manuscript: Four-dimensional quantitative analysis of cell plate development in Arabidopsis using lattice light sheet microscopy identifies robust transition points between growth phases
<p>Representative data accompanying the manuscript: Sinclair R, Wang M, Jawaid MZ, Longkumer T, Aaron J, Rossetti B, Wait E, McDonald K, Cox D, Heddleston J, Wilkop T, Drakakaki G. (2024). <em>Four-dimensional quantitative analysis of cell plate development in Arabidopsis using lattice light sheet microscopy identifies robust transition points between growth phases.</em> J Exp Bot. 2024 Mar 4: erae091. doi: 10.1093/jxb/erae091.</p> <p>The data show YFP–RABA2a dynamics in dividing cells of Arabidopsis root tips using lattice light sheet microscopy. Treatments with or without Endosidin 7, a cytokinesis-specific callose deposition inhibitor, are shown.</p> <p>Data: </p> <p>22.3 YFP-RABA2A. </p> <p>23.9 YFP-RABA2A ES7 </p>
Dataset for the manuscript "Robust ParaHydrogen-Induced Polarization at High Concentrations"
Open the record for dataset details and reuse information.
TIFF file for the frame-wise tracking trajectories in Figure 1 of the manuscript "Assessing the speed of individual bacteria dispersing on mycelial networks"
Open the record for dataset details and reuse information.
fungal-Manuscripts
<p>Data belong to the manuscript “Carbon management measures regulate response of fungal community within soil aggregates to short-term nitrogen and phosphorus additions in temperate forest of northeast China” for submission to the journal Catena.</p> <p> </p> <p>microbial biomass carbon (MBC), microbial biomass nitrogen (MBN), soil exchangeable ammonium (AN) and soil pH (pH), soil organic carbon (SOC), total nitrogen (TN), and total phosphorus (TP).</p>
Data supporting findings for the manuscript: "An Ultra-High Vacuum Scanning Tunneling Microscope with Pulse Tube and Joule-Thomson cooling operating at sub-pm z-noise"
<p>This is the data repository for the manuscript:<br>An Ultra-High Vacuum Scanning Tunneling Microscope with Pulse Tube and Joule-Thomson cooling operating at sub-pm z-noise</p> <p>The data is contained in the zip file.</p> <p>The data is sorted in a folder structure, named after the corresponding images in the manuscript.</p> <p>The raw data and the analysis is given. </p>
Scripts and data for the manuscript "Transcriptomic profiling of gill biopsies to define predictive markers for seawater survival in farmed Atlantic salmon"
<p>This dataset supports the manuscript titled "Transcriptomic profiling of gill biopsies to define predictive markers for seawater survival in farmed Atlantic salmon." It contains comprehensive RNA-seq count data from gill biopsies of approximately 3000 Atlantic salmon smolt, collected during the SynchroSmolt project. The data is supplemented with RNA-seq counts from two prior photoperiod smolt experiments (2013_shortdays and 2017_winterlength) and single-nucleus RNA-seq (snRNA-seq) data from an additional experiment.</p> <p><strong>Key Dataset Elements:</strong><br>- <strong>RNA-seq read counts and metadata</strong> for three experiments, detailing various growth, condition, and survival indicators.<br>- <strong>Scripts for analysis</strong> include differential expression analysis, random forest model preparation and execution, and cell-type-specific gene analysis.<br>- <strong>Intermediate data outputs</strong> such as normalized RNA-seq counts, results from differential expression analyses, and random forest model inputs and outputs.</p> <p><br>This dataset facilitates the exploration of gene expression-based predictive modeling for seawater survival, revealing key insights into the influence of photoperiod history and developmental gene regulation on Atlantic salmon's transition to seawater.</p>
Datasets and relevant code in the Manuscript "Estimation of fire counts and fire radiative power using satellite optical and microwave vegetation indices with random forest method"
<p>1. multiyears_season_fire_ndvi_fwi_edvi_0.25.mat<br>Multiyear averages of ln (FC), ln (FRP), DMC, ISI, EDVI10-18, EDVI18-36, and NDVI over East Asia in 2003–2010</p> <p>2. RF_edvi_data.mat<br>Estimated FC and FRP based on RF model with EDVIs and NDVI </p> <p>3.RF_fwis_data.mat<br>Estimated FC and FRP based on RF model without EDVIs and NDVI </p> <p>4. temporal_variations.mat<br>East Asia Regional Time Series Dataset</p> <p>5.rf_train_cv_forest_review.py<br>Random forest model python code</p>
Raw data for the manuscript: Influence of soil organic matter content on the toxicity of pesticides to soil invertebrates: A review
<p>Files containing Survival (LC50) and reproduction (EC50) data for soil invertebrates exposed to organic chemicals in different soils. The first file contains an overview of all the toxicity data used in the study. The second and third files contain the data used for the "direct comparisons" method, and the fourth and fifth files contain the data (and calculated ratios) used for the "indirect comparisons".</p>
Source data of the MultiSTAAR manuscript "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies".
<p>This dataset serves as the source data for Figures 2-3 and Extended Data Figures 1-2 of the MultiSTAAR manuscript titled "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies". MultiSTAAR is a statistical framework and computationally-scalable analytical pipeline for functionally-informed multi-trait rare variant analysis in large-scale WGS studies.<br><br><strong>Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and ncRNA multi-trait analysis of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerids (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Figure 3.</strong> TOPMed genetic region (2-kb sliding window) unconditional multi-trait analysis results of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerides (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Extended Data Figure 1.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of fasting glucose (FG) and fasting insulin (FI) using TOPMed data (<em>n</em> = 21,731).<br><br><strong>Extended Data Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of C-reactive protein (CRP), interleukin 6 (IL-6), lipoprotein-associated phospholipase A2 (Lp-PLA2) activity, and lipoprotein-associated phospholipase A2 (Lp-PLA2) mass using TOPMed data (<em>n</em> = 9,380).</p>
Raw Data and Analysis Codes for the Dynamic control of the NHSE manuscript
<p>Raw Data and Analysis Codes for the Dynamic control of the NHSE manuscript</p>
Namgyal Manuscript Collection Datasets
<p>These are the official datasets created for the <em>Tibetan Manuscript Project Vienna</em> (<em>TMPV</em>) in the years 2023 and 2024. These datasets contain:</p> <ul> <li>OCR datasets (line image - line label pairs) created from the PageXML annotations</li> <li>PageXML (Transkribus) annotations in Unicode and Wylie</li> <li>PageXML Layout annotations (lines, images, captions, margins) used for image segmentation training</li> <li>OCR models (PyTorch checkpoints and ONNX model files)</li> </ul>
Arid environmental manuscript for wildfire human opinion and burned area data
Open the record for dataset details and reuse information.
Stim circuits for 'Accommodating Fabrication Defects on Floquet Codes with Minimal Hardware Requirements' manuscript
<p>Example Stim circuits for Honeycomb quantum memory experiments with defective qubits.</p> <p>Because accommodating each sample of fabrication defects requires a separate Stim circuit, we only provide example circuits rather than all Stim circuits used in simulations.</p> <p>Please note that the example circuits presented here are constructed by sampling defective qubits according to an iid distribution, with a single parameter defining the probability of an individual qubit being defective. In general a circuit which has a higher probability of each qubit being defective will perform worse than a circuit which has a lower probability of each qubit being defective. However, this does not necessarily mean that the specific circuits presented here will see that behaviour, as these are just individual samples from a large distribution.</p> <p>V2: Uploaded more example circuits with a higher target distance.</p>
Data and code supporting q2-boots manuscript analysis
<div> <div>This record contains data, code, and analysis results for the Raspet et al. q2-boots manuscript. The q2-boots zip file contains the version of the code that was used to perform the analyses presented in the manuscript (commit hash: 1c32499620843de4775ff704d751f6dbd1ce524e). </div> </div>
Code and Data for manuscript "Floral phenotypic divergence and genomic insights in an Ophrys orchid: Unraveling early speciation processes"
<p>Floral phenotypic divergence and genomic insights in an Ophrys orchid: Unraveling early speciation</p> <p>--------</p> <p>This repository contains all the R code used in the manuscript:</p> <p>* Title: "Floral phenotypic divergence and genomic insights in an Ophrys orchid: Unraveling early speciation"</p> <p>* Authors: Anais Gibert, Schatz Bertrand, Buscail Roselyn, Dominique Nguyen, Baguette Michel, Bartes Nicolas and Joris Bertrand</p> <p>* Year of publication: 2024</p> <p>* doi: https://doi.org/10.1101/2024.03.21.586062</p> <p> </p> <p>Synopsis of the study</p> <p>--------</p> <ul> <li> <p>Adaptive radiation in <em>Ophrys</em> orchids leads to complex floral phenotypes that vary in scent, color and shape.</p> </li> <li> <p>Using a novel pipeline to quantify these phenotypes, we investigated trait divergence at early stages of speciation in six populations of <em>Ophrys aveyronensis</em> experiencing recent allopatry. By integrating different genetic/genomic techniques, we investigated: (i) variation and integration of floral components (scent, color and shape), (ii) phenotypes and genomic regions under divergent selection, and (iii) the genomic bases of trait variation.</p> </li> <li> <p>We identified a large genomic island of divergence, associated with phenotypic variation in particular in floral odor. We detected potential divergent selection on macular color, while convergent selection was suspected on floral morphology and for several volatile olfactive compounds. We also identify candidate genes involved in anthocyanin and in steroid biosynthesis pathways associated with standing genetic variation in color and odor.</p> </li> <li> <p>This study sheds light on early differentiation in <em>Ophrys</em>, revealing patterns that often become invisible over time, i.e., the geographic mosaic of traits under selection and the early appearance of strong genomic divergence. It also supports a crucial genomic region for future investigation and highlights the value of a multifaceted approach in unraveling speciation within taxa with large genomes.</p> </li> </ul> <p> </p> <p>Running the code</p> <p>--------</p> <p>Here we present the data and code for carrying out the analyses, as well as the figures and tables from the article and the supplementary material. Once you have installed the necessary packages, run the commands in 'analysis_share.R'. This script uses several functions available in the '/R' directory.</p> <p>Figures and tables are produced in a 'manuscript/figures' and 'manuscript/tables' directory. <br>The `/data' directory contains the data used in the analyses (data/input or data/output), but also the resulting datasets produced by the code (data/RData/).</p> <p> </p>
Pangenomes of multiple species for the "Cluster efficient pangenome graph construction with nf-core/pangenome" manuscript.
<p>Pangenomes of multiple species for the "Cluster efficient pangenome graph construction with nf-core/pangenome" manuscript.</p> <p>Each pangenome is represented in a FASTA format file. Each FASTA file was compressed with <em>bgzip</em> and subsequent indices were created with <em>tabix</em>.</p> <p>The name of each file specifies:</p> <ul> <li>The species,</li> <li>and the number of haplotypes.</li> </ul> <p>The built pangenomes graph are represented in GFA format. Each pangenome graph in the paper is uploaded here, too.</p>
Dataset for manuscript titled "Metal-like ductility and high hardness in nitrogen-rich HfN thin films by point defect superstructuring"
<p>Experimental data for support of claims in the paper. </p>
Dataset for manuscript "Robust Enzyme Discovery and Engineering with Deep Learning using CataPro"
Open the record for dataset details and reuse information.
Data Associated with Manuscript Titled "Evolution of Primate Vocal Repertoires: Vocalization Systems as Embodied Capital for Mediating Within-group Conflict"
<p><span>This is a dataset used in analyses of the macroevolution of primate vocal repertoire size interpreted in the associated manuscript titled "Evolution of Primate Vocal Repertoires: Vocalization Systems as Embodied Capital for Mediating Within-group Conflict." The first tab of the data file contains the following information for each of 42 primates species: maximum longevity (years), endocranial volume (cubic centimeters), log endocranial volume, body mass (g), log body mass, group size, within-group conflict score, vocal repertoire size, and research effort (number of zoological records). The second tab of the data file contains two tables, one reports maximum longevity (years), endocranial volume (cubic centimeters), log endocranial volume, group size, within-group conflict score, and vocal repertoire size values (mean, median, standard deviation, range) aggregated at the suborder, infraorder, superfamily, and family level, while the other reports those statistics aggregated at the family level. The third tab of the data file contains ancestral node ID, ancestral node age (in millions of years), reconstructed ancestral within-group conflict values (mean, 95% lower confidence interval, 95% upper confidence interval), and reconstructed ancestral vocal repertoire size values (mean, 95% lower confidence interval, 95% upper confidence interval). A .pdf file provides visualizations of ancestral character reconstruction (ACR) models with ancestral node IDs for Z-scored within-group conflict (panel A) and Z-scored vocal repertoire size (panel B). </span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.