Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,582
datasets available to search
ShareScore release 0.9.0
Dataset results
1,582 results for “manuscript”
Data and simulation scripts for the manuscript "Patchy charge distribution affects the pH in protein solutions during dialysis."
<p>A zip archive containing the following data:</p> <ul> <li>README.md file describing in detail how to navigate the archive</li> <li>LICENSE.md file describing the license under which the data can be used or reproduced</li> <li>python scripts required to reproduce the plots presented in the manuscript</li> <li>processed simulation results used by the plotting scripts</li> <li>python scripts required to run the simulations in order to reproduce the results</li> </ul> <p> </p> <p> </p>
Dataset and R code support the manuscript titled "Streamlining Linear Free Energy Relationships of Proteins through Dimensionality Analysis and Linear Modeling"
<p>This dataset and R code support the manuscript titled "Streamlining Linear Free Energy Relationships of Proteins through Dimensionality Analysis and Linear Modeling" submitted to the Journal of Chemical Information and Modeling.</p> <p>Table S 1: Chemicals with their experimental values of logKch , and values of logKow and<span> </span>logKaw used to calibrate chicken muscle protein-water 2p-LFER model.</p> <p>Table S 2: Chemicals with their experimental values of logKfish and values of logKow and<span> </span>logKaw used to calibrate fish muscle protein-water 2p-LFER model.</p> <p>Table S 3: Chemicals with their experimental values of logKBSA and values of logKow and<span> </span>logKaw used to calibrate bovine serum albumin-water 2p-LFER model.</p> <p>Table S 4: Chemicals with their experimental values of logKpw and values of logKow and<span> </span>logKaw used to calibrate combined chicken and fish muscle protein-water 2p-LFER model.</p> <p>Table S 5: Diversity of data for logKpw.</p> <p>Table S 6: Diversity of data for logKBSA.</p> <p>Table S 7: Comparison of Experimental and 2p-LFER Predicted Partition Coefficients for ionizable PFAS Compounds.</p> <p>Table S 8: List of neutral fluorotelomer PFAS Compounds.</p> <p>Table S 9: List of experimental in vivo and in vitro partitioning data for different tissues and species.</p> <p>Table S 10: List of experimental Milk-water partition coefficient and predicted values of Milk-water partitioning.</p> <p>Table S 11: Training set for logKpw</p> <p>Table S 12: Validation set for log Kpw</p> <p>Table S 13: Training set for log KBSA</p> <p>Table S 14: Validation set for log KBSA</p>
Project files provided as supporting information to the manuscript "How Communication Pathways Bridge Local and Global Conformations in an IgG4 Antibody: a Molecular Dynamics Study"
<p>June 23, 2021</p> <p>Thomas Tarenzi, Marta Rigoli and Raffaello Potestio</p> <p>==================================</p> <p>The dataset contains the following folders:</p> <p>- contact_area_binding_site: files with the computed surface area, used for the calculation of the contact area between PD-1 and the antibody Fab (Fig. S41).</p> <p>- hbonds_ab-pd1: number of hydrogen bonds between the antigen and the antibody, for each holo cluster (Fig. S41).</p> <p>- mutual_information: matrices with the computed mutual information, for each pair of residues (Fig. S34, S35, S46). The folder contains also the generalized correlation coefficients (Fig. S43) and the correlation scores (Fig. 4, S44, S45), computed from the mutual informations.</p> <p>- networks: communities - for each cluster, each residue is assigned to a community within the interaction network (Fig. S32, S33). betweenness - the values of edge betweenness for each cluster (S30, S31).</p> <p>- output_clustering: each frame of the apo and holo simulations is assigned a cluster index, on the basis of the structural similarity (Fig. S4).</p> <p>- PAD: per-residue values of PAD parameter, for apo and holo systems (Fig. 3, S39).</p> <p>- PCA: principal component analysis for each conformational cluster (Section S2.1).</p> <p>- representative_structures: representative structures for each conformational cluster (Fig. 2).</p> <p>- r_gyr_antibody: radii of gyration of the antibody, for each cluster (Fig. 2).</p> <p>- r_gyr_hinge: radii of gyration of the sole hinge segment, for each cluster (Fig. S36).</p> <p>- RMSD_antibody: distributions of the root-mean-square deviation of the antibody, for each cluster (Fig. S5). </p> <p>- RMSD_antigen: root-mean-square deviation of the antigen PD-1, for each cluster (Fig. S42).</p> <p>- RMSD_binding_site: distributions of the root-mean-square deviation of the residues belonging to the paratope, for each cluster (Fig. S40, S47).</p> <p>- RMSD_matrix: root-mean-square deviation between structures belonging to different pairs of clusters (Fig. S9).</p> <p>- rmsf_antigen: root-mean-square fluctuation of the antigen PD-1, for each cluster (Fig. S42).</p> <p>- rmsf_hinge: difference between the total root-mean-square fluctuations of the two hinge segments, for each cluster (Fig. S38).</p> <p>- salt_bridge: distribution of distances between residues R979 and D1377 (Fig. 4).</p> <p>- sasa_domains: contact area between Fab and Fc antibody domains, for each cluster (Fig. S8).</p> <p>- sasa_hinge: solvent accessible surface area of each hinge segment, for each cluster (Fig. S37).</p>
Figure data for the manuscript "SI-traceable frequency dissemination at 1572.06 nm in a stabilized fiber network with ring topology"
<p>This file contains the data shown in Fig. 1, Fig 3, Fig. 4, Fig. 5 and Fig. 6 of the manuscript "SI-traceable frequency dissemination at 1572.06 nm in a stabilized fiber network with ring topology". Additional information on the data and processing procedure are available from the author upon reasonable request.</p>
Supplementary data for the manuscript: Image2SMILES: Transformer-based Molecular Optical Recognition Engine
<p>This is the supplementary data for the manuscript: <a href="https://chemrxiv.org/engage/chemrxiv/article-details/60c758c6469df4169bf45744">Image2SMILES: Transformer-based Molecular Optical Recognition Engine</a></p> <p>It contains pairs of image-string, generated from 1M SMILES strings. These strings were randomly chosen from PubChem database.<br> It was prepared using the code, published at <a href="https://github.com/syntelly/img2smiles_generator/">https://github.com/syntelly/img2smiles_generator/</a></p> <p>To unpack do:<br> <em>tar xvf subset_1M.tar.xz && tar xvf subset_1M_dump.tar.gz && rm subset_1M_dump.tar.gz</em></p> <p>You'll get the following data:</p> <ul> <li>subset_1M.smi - list of 1M source SMILES</li> <li>subset_1M_dump - directory with images </li> <li>subset_1M_result.csv - list of pairs FGSMILES - pathcode, first 3 chars of pathcode are corresponding subdirs in subset_1M_dump</li> <li>subset_1M_fails.csv - list of failed molecules from subset_1M.smi</li> <li>subset_1M_grpcounter.lst - list of counted groups, used in this generation</li> </ul> <p>You can generate your own data using <a href="https://github.com/syntelly/img2smiles_generator/">https://github.com/syntelly/img2smiles_generator/</a> </p>
Data for the figures in the manuscript about the 2020 pantanal fires
<p>This file contains the data used for the preparation of various figures in the manuscript</p>
Supplemental Material: Page Length Calculations for BMJ EBM Manuscript on COVID-19 Vaccine Transparency
<p>Supplemental Material: Page Length Calculations for COVID-19 Vaccine Transparency Manuscript</p>
Data set related to the manuscript "Mesoscopic simulations of the in situ NMR spectra of porous carbon based supercapacitors: Electronic structure and adsorbent reorganisation effects"
<p>Graphical files in the agr format for all the figures in the manuscript entitled "Mesoscopic simulations of the in situ NMR spectra of porous carbon based supercapacitors: Electronic structure and adsorbent reorganisation effects". Examples of input files for the lattice simulations are also provided.</p>
RNASeq fastq files associated with the manuscript entitled 'Interspecies transcriptome analyses identify genes that control the development and evolution of limb skeletal proportion'
<p>This next-generation sequencing dataset is associated with the research manuscript entitled ‘<em>Interspecies transcriptome analyses identify genes that control the development and evolution of limb skeletal proportion</em>’ (https://www.biorxiv.org/content/10.1101/754002v2).</p> <p>The folder ‘<strong>Zenodo_Saxena_etal_2021_RNASeq_FastqFiles</strong>’ contains raw/unprocessed RNASeq Fatsq read files for postnatal day 5 (P5) mouse (Mus) and jerboa (Jac) cartilage samples (Metatarsal = MT; Radius/Ulna = RU).</p> <p>> The <strong>Jac_P5</strong> subfolder contains single-end reads (R1) for five jerboa metatarsals (MT1-5) and radius/ulna (RU1-5) biological replicates. Jac_MT1-3 and Jac_RU1-3 were used in the primary differential expression analysis (n=3). Jac_MT4-5 and Jac_RU4-5 were used for independent validation (n=2) of the the primary analysis. </p> <p>> The <strong>Mus_P5</strong> subfolder contains single-end reads (R1) for five mouse metatarsals (MT1-5) and radius/ulna (RU1-5) biological replicates. Mus_MT1-3 and Mus_RU1,3 & 4 were used in the primary differential expression analysis (n=3). Mus_MT4-5 and Mus_RU4 & 5 were used for independent validation (n=2) of the the primary analysis. </p>
Data files for manuscript "Exome first approach to reduce diagnostic costs and time – retrospective analysis of 111 individuals with rare neurodevelopmental disorders"
<p>#2021-07-23<br> #Summary<br> This ZIP-file contains the Excel files used for the clinical and variant analyses for the manuscript "Exome first approach to reduce diagnostic costs and time – retrospective analysis of 111 individuals with rare neurodevelopmental disorders".</p> <p>#Folder structure<br> ./ (parent directory containing this README file and all subfolders)<br> ./Clinical/ (contains an Excel sheets with complete clinical data, costs and criteria)<br> ./Variants/ (contains an Excel sheet with all variant annotation)</p> <p>#Files and checksums<br> 6DB0EF10BE7A7AF5A18E523F33FB662A ./Clinical/FileS2_Clinical.xlsx<br> 0B0C1AFA0751B63A35DC99DB25546A38 ./Variants/FileS3_Variants.xlsx</p>
W-band dataset with I/Q measurement for an AMT manuscript
<p>The dataset contains 3 min of I/Q measurements from the W-band radar operated by TU Delft in Cabauw, the Netherlands. The data file has a binary format. The dataset was used in a publication which is currently in preparation. The archive contains the pdf file (IQ_Format.pdf) with the description of the data file content.</p>
Data and Code for the manuscript PhD students in life sciences can benefit from team cohesion
<p>This folder contains data and code to reproduce regression results of the manuscript "PhD students in life sciences can benefit from team cohesion".</p>
Grombacher et al. surface NMR data GRL manuscript
<p>Data supporting a manuscript submitted to Geophysrical research letters.</p>
Accompanying data set for the manuscript "REverSe TRanscrIptase Chain Termination (RESTRICT) for Selective Measurement of Nucleotide Analogs Used in HIV Care and Prevention"
<p>This data set contains all experimental and theoretical data included in the manuscript " REverSe TRanscrIptase Chain Termination (RESTRICT) for Selective Measurement of Nucleotide Analogs Used in HIV Care and Prevention", namely:</p> <p>RESTRICT_model: MATLAB script for completing calculations in the RESTRICT theoretical model.</p> <p>Figure 2:</p> <ul> <li>Raw data from theoretical model showing contributions of individual model components, Kaff = 0.3</li> <li>Normalized data from theoretical model showing contributions of individual model components, Kaff = 0.3</li> <li>Experimental NRTI Drug Screen 180 nt TTCA 500 nM dNTP</li> </ul> <p>Figure 3:</p> <ul> <li>Experiment-dNTP-Concentration-Screen</li> <li>Theory-dNTP-Concentration-Screen</li> <li>Experiment-Template-Length-Screen</li> <li>Theory-Template-Length-Screen</li> <li>Experiment-Sequence-Screen</li> <li>Theory-Sequence-Screen</li> </ul> <p>Figure 4:</p> <ul> <li>Experimental-NRTI-Drug-Screen-90nt-TCAA-only</li> <li>Theory-TCAA90-Kaff=0point2</li> <li>Experimental-NRTI-Drug-Screen-GGCA-only</li> <li>Theory-GGCA180-Kaff=0point2</li> </ul> <p>Figure 5:</p> <ul> <li>GGCA-vs-TTCA-Specificity-Analysis</li> </ul>
Dataset for manuscript entitled "The effects of a synthetic and biological surfactant on the community composition and metabolic activity of a freshwater biofilm"
<p>The following datasets were used for the 16s rRNA analysis in the manuscript entitled " The effects of a synthetic and biological surfactant on the community composition and metabolic activity of a freshwater biofilm". BZ2 files were obtained from next generation sequencing with the Illumina Mi-Seq. Mothur was used to analyze the BZ2 files, creating the listed excel documents.</p>
Data for manuscript "Prevalence in News Media of two Competing Hypotheses about COVID-19 Origins"
<p>The Covid-19 pandemic has been one of the most disruptive and painful phenomena of the last few decades. As of July 2021, the origins of the SARS-CoV-2 virus that caused the outbreak remain a mystery. This work analyzes the prevalence in news media articles of two popular hypotheses about SARS-CoV-2 virus origins: the natural emergence and the lab-leak hypotheses. </p> <p>This data set contains frequency counts of target words in news and opinion articles from 12 popular news media outlets. The target words are listed in the associated manuscript and are mostly words associated with the Covid-19 pandemic. </p> <p>The list of compressed files in this data set is listed next:</p> <p>targetWordsInArticlesCounts.rar contains counts of target words in outlets articles as well as total counts of words in articles</p> <p>targetWordsFrequencies.rar daily, weekly, monthly word frequencies</p> <p>wordEmbeddingModels.rar monthly embedding models of news outlets content</p> <p>analysisScripts.rar analysis notebooks</p> <p>The textual content of news and opinion articles from the outlets is available in the outlet's online domains and/or public cache repositories such as Google cache, The Internet Wayback Machine, and Common Crawl. We used derived word frequency counts from these sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific XPath expressions. Tokens were lowercased prior to estimating frequency counts. </p> <p>Yearly frequency usage of a target word in an outlet in any given temporal interval ( daily, weekly, monthly) was estimated by dividing the total number of occurrences of the target word in all articles of a given temporal interval by the number of all words in all articles of that temporal interval. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>In a small percentage of articles, outlet specific XPath expressions might fail to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles are not precise. In a random sample of articles and outlets, manual estimation of target words counts overlapped with the automatically derived counts for over 90% of the articles. Most of the incorrect frequency counts are minor deviations from the actual counts such as for instance counting a word in an article footnote encouraging article readers to find related articles and that the XPath expression might mistakenly include as the content of the article main text. Some additional outlet-specific inaccuracies that we could identify occurred in the WSJ where in less than 5% of the articles XPath expressions failed to capture the article's main text content. Other outlets articles samples sizes might not be comprehensive but, to the best of our knowledge, they are representative and include tens of thousands of articles per outlet/year. To conclude, in a data analysis of over 1.5 million articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 1 of main manuscript for supporting evidence).</p> <p> </p> <p> </p>
A supplementary file for manuscript submission - Videos of flume test Events
<p>The ZIP file contains 7 videos clips of a flume test for landslide dam breach. The videos are the Supplementary Materials for a manuscript to be submitted to a journal in August, 2021 for possible publication. The ZIP file is uploaded on August 19, 2021 by Prof. Zheng-yi Feng.</p>
Data manuscript: Preventive training does not interfere with mRNA-encoding myosin and collagen expression during pulmonary arterial hypertension
<p>Supporting Information files of manuscript: Preventive training does not interfere with mRNA-encoding myosin and collagen expression during pulmonary arterial hypertension </p> <p> </p> <p>The values behind the means, standard deviations and values used to build graphs; </p>
Online Package for the manuscript "Continuous Integration and Delivery for Cyber-Physical systems: A Grounded-Theory"
<p>This package contains the material of the (grounded theory) study related to the paper</p> <p>"Continuous Integration and Delivery for Cyber-Physical systems: A Grounded-Theory"</p> <p>The content of the various files and directories is detailed in the following.</p> <p>InterviewStructure.pdf: complete interview structure (the paper reports an overview in Table 1)</p> <p>open_coding/: this directory contains details from the open coding procedure. In particular:<br> - AllCodesUsedWhileCodingAndMapping.csv contains the list of all codes generated during the open coding, with a symbol near each one (first column) then used to compute the inter-rater agreement<br> - first_round.csv, second_round.csv, third_round.csv, fourth_round.csv: files used to compute the inter_rater agreement</p> <p><br> CodesContributingToMindMap.csv: final set of codes, that contributed to the taxonomy (see below)</p> <p>FinalCodingTraceability.csv: this file describes how the final set of codes is traced onto the ten interviews. Note that, at this stage, the transcripts have been redacted for confidentiality purposes.</p> <p>MindMap_Complete.pdf: complete taxonomy of codes, in the form of a mind map. The one reported in the paper (Figure 2) cuts leaves, unless (as for benefits, for example) they are necessary to properly describe and understand the category. Also, note that the complete mind map separates the benefits into "actual" and "expected" (based on what was collected from the interviews) whereas the summary mindmap shown in the paper (Figure 2) does not make this distinction.</p> <p>6C_Diagrams/ : This directory contains the detailed 6C diagrams (i.e., each box related to a "C" contains the list of codes pertaining to it) for the three dimensions investigated in the paper and addressed in RQ1, RQ2, and RQ3.</p>
Data and code for the manuscript: "An ecological explanation for hyperallometric scaling of reproduction"
<p>Code, data, and data objects underlying the analyses presented in the manuscript "An ecological explanation for hyperallometric scaling of reproduction".</p> <p>A version of the manuscript is available on bioRxiv: https://doi.org/10.1101/2021.03.12.435090</p> <p>Abstract for the manuscript:</p> <p>"In wild populations, large individuals have disproportionately higher reproductive output than smaller individuals. Some theoretical models explain this pattern – termed reproductive hyperallometry – by individuals allocating a greater fraction of available energy towards reproductive effort as they grow. Here, we propose an ecological explanation for this observation: differences between individuals in rates of resource assimilation, where greater assimilation causes both increased reproduction and body size, resulting in reproductive hyperallometry at the level of the population. We illustrate this effect by determining the relationship between size and reproduction in wild and lab-reared Trinidadian guppies. We show that (i) reproduction increased disproportionately with body size in the wild but not in the lab, where resource competition was eliminated; (ii) in the wild, hyperallometry was greatest during the wet season, when resource competition is strongest; and (iii) detection of hyperallometric scaling of reproduction at the population level was inevitable if individual differences in assimilation were ignored. We propose that ecologically-driven variation in assimilation – caused by size-dependent resource competition, niche expansion, and chance – contributes substantially to hyperallometric scaling of reproduction in natural populations. We recommend that mechanistic models incorporate such ecologically-caused variation when seeking to explain reproductive hyperallometry.'</p> <p>The zip file contains:</p> <p>1. A single R script for running the analyses, generating the results, and producing the figures given in the manuscript </p> <p>- size_reproduction_scaling_lab_wild.R </p> <p>2. Two .csv files containing the wild and lab datasets used in the study:</p> <p>- wild_guppy_data_81-87.csv (the wild dataset)<br> - senecence_long_data.csv (the lab dataset)</p> <p>3. A folder containing 14 model objects. These are the models produced by running the R script. These are provided to save the user time. The model object names correspond to the models described in Table 1 of the manuscript.</p> <p>4. A READ_ME.txt file.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.