Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,582
datasets available to search
ShareScore release 0.9.0
Dataset results
1,582 results for “manuscript”
Raw and analyzed data for manuscript "Reduction of copper surface oxide using a sub-atmospheric dielectric barrier discharge plasma"
<p><strong>Abstract:</strong> Oxide layers on metal surfaces adversely affect processability and material properties in many industrial applications. Although several plasma-based approaches for deoxidation were investigated in the past, oftentimes they either work under conditions expensive to create or require long processing times. In this study, the deoxidation effect of a non-thermal dielectric barrier discharge plasma in an Ar/H<sub>2</sub> gas mixture at 100 hPa and 20 °C was investigated on copper surfaces with a native oxide layer. The chemical structure of surfaces before and after deoxidation was analyzed by X-ray photoelectron spectroscopy (XPS). The results revealed that ~98 % of the surface lattice oxide Cu<sub>2</sub>O was reduced to Cu after around 20 s of plasma treatment, whereas all oxygen contaminants were almost completely removed from Cu surface after around 50 s. Additionally, the kinetics of the reduction of surface oxide was studied and a Johnson-Mehl-Avrami-Erofeev-Kholmogorov kinetic model was proposed. The analysis of the morphology of surfaces was performed with atomic force microscopy (AFM), showing minor changes in the roughness after deoxidation. Moreover, optical emission spectroscopy (OES) showed atomic hydrogen radicals in the plasma phase, which likely causes deoxidation effect.</p>
Data for the manuscript entitled "Factors driving large-scale ungulate carrion production in the Anthropocene"
<p>Data for the article entitled "Factors driving large-scale ungulate carrion production in the Anthropocene". The data includes 5 main datasets, each of those corresponding to 5 main carrion production sources in terrestrial ecosystems in peninsular Spain, namely; 1) Livestock, 2) Big Game Hunting, 3) Roadkills, 4) Predation and 5) Natural mortality. </p> <p>In case of any doubt/s or enquiries regarding this data, please, send an email to the corresponding author; Jon Morant Etxebarria (email: jmorant@aranzadi.eus). </p> <p> </p>
Datasets for the manuscript "In silico proof of principle of machine learning-based antibody design at unconstrained scale"
<p>The zip file contains dataset files for the manuscript "In silico proof of principle of machine learning-based antibody design at unconstrained scale"</p>
Data files for manuscript "Prenatal phenotype of PNKP-related primary microcephaly associated with variants in the FHA and Phosphatase domain"
<p>#2021-08-26<br> #Summary<br> This ZIP-file contains the Excel files used for the clinical and variant analyses of the PNKP protein/gene for the manuscript "Prenatal phenotype of PNKP-related primary microcephaly associated with variants in the FHA and Phosphatase domain".</p> <p>#Folder structure<br> ./ (parent directory containing this README file and all subfolders)<br> ./Clinical/ (contains an Excel sheet with complete clinical data of 4 affected individuals)<br> ./Variants/ (contains an Excel sheet with all variant annotation and domain information used in 6 sheets)</p> <p>#Files and checksums<br> 4B5BE914E77EB6511DF3DB94C6918E45 ./Clinical/FileS2_PNKP_Clinical.xlsx<br> 0228C6EF543D597C1C800B8D78E3097D ./Variants/FileS3_PNKP_Variants.xlsx</p>
Dataset for submitted manuscript "Temperatures and cooling rates recorded by the New Caledonia ophiolite: implications for cooling mechanisms in young forearc sequence"
<p>original dataset for the manuscript "Temperatures and cooling rates recorded by the New Caledonia ophiolite: implications for cooling mechanisms in young forearc sequences" submitted to the journal "G3 Geophysics, Geochemistry, Geosystems"</p>
Data files for manuscript "Prevalence of hereditary tubulointerstitial kidney diseases in the German Chronic Kidney Disease study"
<p>#2021-09-19<br> #Summary<br> This ZIP-file contains the Excel files used for all analyses for the manuscript "Prevalence of hereditary tubulointerstitial kidney diseases in the German Chronic Kidney Disease study".</p> <p><br> #File structure<br> README.txt This README file.<br> File S1 ("FileS1_GCKD-ADTKD.cohort.xlsx") Cohort characteristics, sequencing quality parameters and fingerprinting results.<br> File S2 ("FileS2_GCKD-ADTKD.content.xlsx") Sequencing panel design/ content with information on gene domains used for Figure 2.<br> File S3 ("FileS3_GCKD-ADTKD.variants.xlsx") Information on small variants, CNVs and MUC1 analyses (SNaPshot and adVNTR).<br> File S4 ("FileS4_GCKD-ADTKD.simulation.xlsx") Curated variant and individual data from the Groopman study with results of the simulation for Figure 4.</p> <p><br> #Files and checksums<br> 2D6184BC145987D3EE2E0DC6873DDDDB ./README.txt<br> 10EE4C3C9D9ED6417647F26A09583820 ./FileS1_GCKD-ADTKD.cohort.xlsx<br> 9F95661845215FB7875EB053F887AB36 ./FileS2_GCKD-ADTKD.content.xlsx<br> 2AC19FBC6103CB48905F64716428AE5B ./FileS3_GCKD-ADTKD.variants.xlsx<br> 2D6184BC145987D3EE2E0DC6873DDDDB ./FileS4_GCKD-ADTKD.simulation.xlsx</p>
Supplementary data for the manuscript "Comparison of computational and experimental saturation vapor pressures of α-pinene + O3 oxidation products"
<p>COSMO-files of potential ozonolysis products of α-pinene for the manuscript:<br> Hyttinen, N., Pullinen, I., Nissinen, A., Schobesberger, S., Virtanen, A., and Yli-Juuti, T.: Comparison of computational and experimental saturation vapor pressures of α-pinene + O<sub>3</sub> oxidation products, Atmos. Chem. Phys. Discuss. [preprint], https://doi.org/10.5194/acp-2021-775, in review, 2021.</p>
Faithful Transcriptions Data Set: TEI/XML-encoded Transcriptions of Medieval Theological Manuscripts
<p>From May to July 2021, the Berlin State Library and the Leipzig University Library jointly organized the Transcribathon <a href="https://lab.sbb.berlin/events/faithful-transcriptions/">“Faithful Transcriptions”</a>, a digital crowd souring project on medieval theological manuscripts. During the project, over 100 participants produced TEI/XML-encoded transcriptions in the IIIF workspace of the <a href="https://handschriftenportal.de/">Handschriftenportal</a>, which is currently being developed.</p> <p>The <a href="https://lab.sbb.berlin/datensets-transkribathon/">Faithful Transcriptions Data Set</a> contains 181 pages with 8.952 text lines from 12 manuscripts in German, Dutch, and Latin. The medieval scripts include Textura, Textualis, Gothic Cursiva, and Bastarda. The transcriptions are linked to the coordinates of the digitized manuscript image on text line level.</p> <p>--------------------------------------------</p> <p>Von Mai bis Juli 2021 richtete die Staatsbibliothek zu Berlin in Kooperation mit der Universitätsbibliothek Leipzig den Transkribathon <a href="https://lab.sbb.berlin/events/faithful-transcriptions/">„Faithful Transcriptions“</a> aus, ein digitales Crowd-Sourcing-Projekt zu theologischen Handschriften des Mittelalters. Über 100 Teilnehmende fertigten dabei TEI/XML-codierte Transkriptionen in der IIIF-basierten Arbeitsumgebung des aktuell in Entwicklung befindlichen <a href="https://handschriftenportal.de/">Handschriftenportals</a> an. </p> <p>Das <a href="https://lab.sbb.berlin/datensets-transkribathon/">Faithful Transcriptions-Datenset</a> enthält 181 Seiten mit 8.952 Textzeilen aus 12 Handschriften in deutscher, niederländischer und lateinischer Sprache. Die mittelalterlichen Schriften reichen von Textura über Textualis und Gotische Kursive bis hin zur Bastarda. Die Transkriptionen sind mit den Bildkoordinaten des Handschriftendigitalisats auf Textzeilenebene verknüpft. </p>
Supplement Data to Manuscript "Allocating small transporters to large jobs"
<p><strong>Data Description</strong></p> <p>In the data file `instances.csv`, we provide the original data that was used for the manuscript `Allocating small transporters to large jobs` by Neil Jami, Neele Leithäuser and Christian Weiß from Fraunhofer ITWM. </p> <ul> <li>SimulationId: Identifier for each configuration (encoded by NumberJobs_NumberResources).</li> <li>InstanceIndex: Running index for each sample of a simulation configuration.</li> <li>NumberJobs: Number of jobs in the simulation configuration.</li> <li>NumberResources: Number of available resources (transporters) in the simulation configuration.</li> <li>JobReturnTimes: Array of return times for the individual jobs in this sample.</li> <li>NumberResourceWithCapacity_1: Number of resources in this sample that have Filling time =1.</li> <li>NumberResourceWithCapacity_2: Number of resources in this sample that have Filling time =2.</li> <li>NumberResourceWithCapacity_3: Number of resources in this sample that have Filling time =3.</li> <li>NumberResourceWithCapacity_4: Number of resources in this sample that have Filling time =4.</li> <li>NumberResourceWithCapacity_5: Number of resources in this sample that have Filling time =5.</li> </ul>
Mp4-Version of the supplementary Movies for the manuscript: "Quantitative real-time in-cell imaging reveals heterogeneous clusters of proteins prior to condensation"
<p>Videos in 'mp4'-format of the 8 supplementary movies for the manuscript: "Quantitative real-time in-cell imaging reveals heterogeneous clusters of proteins prior to condensation"</p>
Intermediate files accompanying manuscript "The genomic basis of reproductive and migratory behaviour in a polymorphic salmonid".
<p>Recent ecotypic differentiation provides unique opportunities to investigate the genomic basis and architecture of local adaptation, while offering insights into how species form and persist. Sockeye salmon (<em>Oncorhynchus nerka</em>) exhibit migratory and resident (‘kokanee’) ecotypes, which are further distinguished into shore-spawning and stream-spawning reproductive ecotypes. Here, we analysed 36 sockeye (stream-spawning) and kokanee (stream- and shore-spawning) genomes from a system where they co-occur and have a recent common ancestry (Okanagan Lake/River in British Columbia, Canada) to investigate the genomic basis of reproductive and migratory behaviour. Examination of the genomic landscape of differentiation, differences in allele frequencies, and genotype-phenotype associations revealed three main blocks of sequence differentiation on chromosomes 7, 12, and 20, associated with migratory behaviour, spawning location, and spawning timing. Several structural variants identified in these same areas suggest they could contribute to ecotypic differentiation directly as causal variants or via maintenance of their genomic architecture through recombination suppression mechanisms. Genes in these regions were related to spatial memory and swimming endurance (<em>SYNGAP</em>, <em>TPM3</em>), as well as eye and brain development (including <em>SIX6</em>), potentially associated with differences in migratory behaviour and visual habitats across spawning locations, respectively. Additional genes (<em>GREB1L</em>, <em>ROCK1</em>) identified have been associated with timing of migration in other salmonids and could explain variation in timing of spawning in our study. Together, these results based on the joint analysis of sequence and structural variation represent a significant advance in our understanding of the genomic landscape of ecotypic differentiation at different stages in the speciation continuum.</p>
Raw data files for the manuscript entitled "Understanding your support system: the design of a stable MOF/polyazoamine support for biomass conversion"
<p>Raw data files for the manuscript entitled "Understanding your support system: the design of a stable MOF/polyazoamine support for biomass conversion" published in Chemistry of Materials. <a href="https://doi.org/10.1021/acs.chemmater.2c01731">https://doi.org/10.1021/acs.chemmater.2c01731</a></p> <p>The files are organized by manuscript figure names and are in a simple text format. The headers contain the necessary information such as column designations.</p>
Dataset presented in the recently submitted AGU manuscript "Constraining the crustal and mantle conductivity structures beneath islands by a joint inversion of multi-source magnetic transfer functions"
<p>Dataset (observed tippers, solar quiet global-to-local transfer functions, and global Q responses) presented in the recently submitted AGU manuscript "Constraining the crustal and mantle conductivity structures beneath islands by a joint inversion of multi-source magnetic transfer functions".</p>
Comparison of Large Eddy Simulations against measurements from the Lillgrund offshore wind farm - Manuscript data
<p>Time averaged power and farm inflow velocity for the manuscript "Comparison of Large Eddy Simulations against measurements from the Lillgrund offshore wind farm" for publication in the wind energy science journal. Data is uploaded for the 5 simulation cases covered.</p> <p>'Power' files contain average power production for 48 turbines. First row corresponds to LES data, second row corresponds to SCADA data from the Lillgrund wind farm.</p> <p>'Velocity' files contain inflow mean velocity measurements at the 72 range gate locations. First row corresponds to LES inflow data, second row corresponds to LIDAR inflow data from the Lillgrund wind farm.</p>
cazy_webscraper manuscript additional files
<p>Additional files that accompany the manuscript: cazy_webscraper: local compilation and interrogation of comprehensive CAZyme datasets</p>
KMCP Manuscript Data
<p># KMCP: accurate metagenomic profiling of both prokaryotic and viral populations by pseudo-mapping</p> <p>## 1.code-and-documents</p> <p>This directory contains the source code, executable binaries, and documents of KMCP,<br> which are also hosted at Github: https://github.com/shenwei356/kmcp .</p> <p>Databases, usage, and tutorials of KMCP are also available at https://bioinf.shenwei.me/kmcp/.</p> <p>- [Installation](https://bioinf.shenwei.me/kmcp/download)<br> - [Databases](https://bioinf.shenwei.me/kmcp/database)<br> - Tutorials<br> - [Taxonomic profiling](https://bioinf.shenwei.me/kmcp/tutorial/profiling)<br> - [Sequence and genome searching](https://bioinf.shenwei.me/kmcp/tutorial/searching)<br> - [Usage](https://bioinf.shenwei.me/kmcp/usage)<br> - [Benchmarks](https://bioinf.shenwei.me/kmcp/benchmark)<br> - [FAQs](https://bioinf.shenwei.me/kmcp/faq)</p> <p>## 2.databases</p> <p>This directory contains the building steps and reference genome accessions for<br> KMCP databases used in the manuscript.</p> <p> cami2 Databases used in benchmarks on CAMI2 mouse gut datasets<br> kmcp Databases used in other benchmarks<br> <br> ## 3.figures</p> <p>Each subdirectory contains steps to run the benchmark (`README.md`), steps for plotting (`README-plot.md`),<br> benchmark results, and figures.<br> </p>
Data and materials for "The Consequences of Data Dispersion in Genomics: A Comparative Analysis of Data Sources for Precision Medicine" manuscript"
<p>Data and sripts for the "The Consequences of Data Dispersion in Genomics: A Comparative Analysis of Data Sources for Precision Medicine" manuscript" manuscript, sent to BMC Bioinformatics</p>
Immunostaining data for reviewing manuscript
<p>The reviewing manuscript entitled "CRISPR/Cas9 Mediated Specific Ablation of Vegfa in Retinal Pigment Epithelium Efficiently Regresses Choroidal Neovascularization" show immunostaining in the figure plates. The entire immunostaining images series are provided here as .tif image sequences. The file names refer to the figure numbers and position in the figure plates.</p>
Datasets of the manuscript "Rational design of profile HMMs for sensitive and specific sequence detection with case studies applied to viruses, bacteriophages, and casposons"
<p><strong>DATASETS</strong></p> <p>Rational design of profile HMMs for sensitive and specific sequence detection with case studies applied to viruses, bacteriophages, and casposons</p> <p>Liliane S. Oliveira, Alejandro Reyes, Bas E. Dutilh and Arthur Gruber<sup>*</sup></p> <p>* Correspondence: <a href="mailto:argruber@usp.br">argruber@usp.br</a> (AG); Tel. +55 11 3091 7274</p> <p> </p> <p>Here we provide different data of <em>Microviridae</em>, <em>Flavivirus</em> and casposons used throughout the work:</p> <ul> <li>Microviridae folder <ul> <li>conserved_HMMs – profile HMMs constructed with TABAJARA in Conservation mode for <em>Microviridae</em></li> <li>discriminative_HMMs – profile HMMs constructed with TABAJARA in Discrimination mode for <em>Microviridae</em></li> <li>sequences – different sequence datasets and respective multiple sequence alignments <ul> <li>Microviridae_113-seq_training_set.fasta - 113 VP1 sequences covering diversity of the <em>Microviridae</em> family</li> <li>Microviridae_113-seq.aln – multiple sequence alignment of the 113-protein dataset</li> <li>Microviridae_1836-seq_testset.fasta - 1,836 sequence dataset covering 1,836 sequences of the major capsid protein (VP1) comprising 501 <em>Alpavirinae</em> sequences, 1,040 <em>Gokushovirinae</em> sequences and 295 <em>Pichovirinae</em> sequences</li> <li>Microviridae_1866-seq.aln - multiple sequence alignment of the 1,866-protein <em>Microviridae</em> dataset used in the experiment of Figure 4</li> </ul> </li> </ul> </li> <li>Flavivirus folder <ul> <li>conserved_HMMs – profile HMMs constructed with TABAJARA in Conservation mode for <em>Flavivirus</em></li> <li>discriminative_HMMs – profile HMMs constructed with TABAJARA in Discrimination mode for <em>Flavivirus</em> <ul> <li>full-length – models constructed from full-length protein sequences</li> <li>short - models constructed from selected short alignment blocks of the protein sequences</li> </ul> </li> <li>sequences – different sequence datasets and respective multiple sequence alignments <ul> <li>Flavivirus_127-seq_training_set.fasta - 127 polyprotein sequences covering species diversity of the genus <em>Flavivirus</em></li> <li>Flavivirus_127-seq.aln – multiple sequence alignment of the 127-protein dataset</li> <li>Flavivirus_6364-seq_testset.fasta - 6,364 sequence dataset covering species diversity of <em>Flavivirus</em>, including 3,919 of dengue virus (DENV), 327 of Zika virus (ZIKV), 63 of yellow fever virus (YFV), and the remaining 2,055 sequences covering other available flaviviruses</li> <li>Flavivirus_6364-seq.aln - multiple sequence alignment of the 6,364-protein <em>Flavivirus</em> dataset</li> </ul> </li> </ul> </li> <li>Casposons folder <ul> <li>casposon_generic_HMMs – profile HMMs constructed with TABAJARA in Discrimination mode for the generic detection of all casposons and discrimination from CRISPRs.</li> <li>casposon_family_discriminative_HMMs – profile HMMs constructed with TABAJARA in Discrimination mode for the specific discrimination among casposon families and from CRISPRs.</li> <li>sequences – different sequence datasets and respective multiple sequence alignments <ul> <li>casposons_crisprs.fasta – 106 Cas1 <em>bona fide</em> sequences derived from 52 CRISPRs and 54 casposons</li> <li>casposon_family_discrimination.aln - multiple sequence alignment of 52 <em>bona fide</em> CRISPR and 54 casposon sequences, with appropriate nomenclature to run TABAJARA for the discrimination of each casposon family.</li> <li>casposons_crisprs_discrimination.aln - multiple sequence alignment of 52 <em>bona fide</em> CRISPR and 54 casposon sequences, with appropriate nomenclature to run TABAJARA for discrimination of CRISPRs and casposons.</li> </ul> </li> </ul> </li> </ul>
Dataset and evaluation for HTR models for Latin and French Medieval Documentary Manuscripts
<p><strong>1. Dataset presentation.</strong></p> <p>This is the dataset used to produce the HTR models applied to documentary Latin and French manuscripts presented in the paper: Sergio Torres Aguilar, Vincent Jolivet. <strong>Handwritten Text Recognition for Documentary Medieval<br> Manuscripts. </strong>2022. https://hal.science/hal-03892163</p> <p>The dataset contains mostly charters and registers from the Late-medieval period (12th-15th). The training and evaluation, entailing 1855 pages, 120k lines of text and almost 1M tokens, were conducted using three freely available ground-truth corpora :</p> <p><strong>The Alcar-HOME database </strong>: https://zenodo.org/record/5600884</p> <p><strong>The e-NDP corpus </strong>: https://github.com/chartes/e-NDP_HTR</p> <p><strong>The Himanis project </strong>: https://zenodo.org/record/5535306</p> <p>The final model operates in a multilingual environment (Latin and French) and it is able to recognize several Latin script families (mostly <em>Textualis</em> and <em>Cursiva</em>) in documents produced in ca. 12th - 15th centuries. During the evaluation the models shows an accuracy of <strong>94.01%</strong> on the validation set and a CER (character error ratio) of about <strong>0.12</strong> to <strong>0.17</strong> on four external unseen datasets. A fine-tuning exercise using 10 ground-truth pages can raise these results to a CER between <strong>0.06</strong> to <strong>0.10</strong> respectively.</p> <p> </p> <p><strong>2. Dataset contents .</strong></p> <p>a) <em>GT_list : </em>List containing the GT file names which constitute the training, evaluation and test sets. The images and transcriptions can be downloaded from their original repositories.</p> <p>b) <em>Training :</em> Contains the training and testing results (evaluation and prediction files) presented in the original paper for the two training phases: Regular (Textualis and Cursiva separated training) and Quartiles (mixed training by quartiles).</p> <p>c) <em>Useful_scripts :</em> Scripts to produce the HTR metrics (CER, WER, SER) and plot the model's accuracy.</p> <p>d) <em>Best_model :</em> Contains the best multilingual and multi-script model.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.