Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
10,812
datasets available to search
ShareScore release 0.9.0
Dataset results
10,812 results for “novels”
SpatialMETA: A Novel Framework for Integrating Spatial Transcriptomics and Metabolomics Data
<p>Multimodal analysis of spatial transcriptomics (ST) and spatial metabolomics (SM) has rapidly advanced for characterizing tissue microenvironments. However, integrating ST and SM data remains challenging due to differing morphologies, resolutions, and batch effects. We developed SpatialMETA (Spatial Metabolomics and Transcriptomics Analysis), a novel method for integrating spatial multi-omics data, which aligns ST and SM to a unified resolution, enables both cross-modal and cross-sample integration to identify ST-SM associated spatial patterns, and provides extensive visualization and analysis capabilities. The datasets for SpatialMETA is avaiable. </p>
Coral thermotolerance retained following year-long exposure to a novel environment
<p>Description of datasets:</p> <h4>HTSeqCounts</h4> <p>Raw HTSeq count data used for downstream RNA sequencing analyses and visualisations including Principal Component Analysis and Differential Gene Expression analysis of the four habitat groups (mangrove to reef, reef to reef, wild mangrove and wild reef).</p> <h4>GO enriched gene results:</h4> <p>GO enriched genes summarised by broad categories of GO Slims for <em>Pocillopora acuta</em> corals originating from a mangrove system vs colonies from an adjacent reef site, translocated to a reef environment for one year. Samples for gene expression were collected during an acute heat stress assay to assess differentially enriched genes between mangrove and reef corals under 36 ºC. Enriched GO terms were grouped into broader functional categories using the “goSlim” function in the GSEABase R package with GOslim generic obo as the reference database to allow for a summary of key functions enriched in groups under heat stress. See manuscript methods for further detail on differential gene expression analysis.</p> <h4>2022-2023 mangrove/reef temperature and pH</h4> <p>Temperature and pH data measured in the Low Isles mangrove lagoon and reef habitat using HOBO MX2510 loggers deployed from February 2022 - February 2023. </p> <h4>Methylated DNA data</h4> <p>Percent DNA methylation relative to total DNA of mangrove to reef, reef to reef and wild mangrove groups under acute temperature stress during the February 2023 CBASS experiment.</p> <h4>Analysis code.zip</h4> <p>A copy of all R code used in the data analysis of this project</p> <p> </p>
Genome-wide association analyses identify novel Brugada syndrome risk loci and highlight a new mechanism of sodium channel regulation in disease susceptibility
<p>The Brugada syndrome GWAS summary statistics</p> <p>Brugada syndrome is a cardiac arrhythmia disorder associated with sudden death in young adults. With the exception of <em>SCN5A</em>, encoding the cardiac sodium channel Na<sub>V</sub>1.5, susceptibility genes remain largely unknown. We performed a genome-wide association meta-analysis comprising 2,820 unrelated cases with Brugada syndrome and 10,001 controls.</p> <p> </p>
Simulated NGS read datasets for prediction of novel fungal pathogens and multiple pathogen classes
<p>This repository contains simulated Illumina read datasets for novel fungal pathogen prediction and real-time detection of multiple pathogen classes. They were used to train the models hosted at <a href="https://zenodo.org/record/5711877">https://zenodo.org/record/5711877</a>.<br> The reads were simulated with Mason (<a href="https://www.seqan.de/apps/mason/">https://www.seqan.de/apps/mason/</a>) from genomes downloaded from NCBI, based on metadata stored in a manually curated database (<a href="https://zenodo.org/record/5846345">https://zenodo.org/record/5846345</a>).</p> <p>We provide the following:</p> <p>1) An rds file describing assignment of fungal species from the database (<a href="https://zenodo.org/record/5846345">https://zenodo.org/record/5846345</a>) to training, validation and test sets (TrainValTest_fungi.rds). A second rds file (TrainValTest_temporal.rds) includes species added within 12 weeks after the original datasets were compiled. Those species were used for a temporal benchmark.</p> <p>2) Fungal validation and test sets. Each contains 1.25 million, 250bp-long reads simulated from non-overlapping sets of human ("pathogenic") or non-human ("nonpathogenic") pathogens. The test set contains paired reads ("_1" and "_2" for the first and second mate). The number of reads per species is proportional to the respective genome length. An additional, temporal test set (*temporal*fasta.gz) includes 15 species added after 12 weeks from the consturction of the original datasets.</p> <p>3) Fungal training sets. They contain 250bp-long reads simulated from species not present in the validation or test sets. There are four variants:<br> 3a) "low-coverage, linear" - 20 million reads, number of reads per species proportional to genome length<br> 3b) "low-coverage, logarithmic" - 20 million reads, number of reads per species proportional to the logarithm of genome length ("log")<br> 3c) "high-coverage, linear" - 240 million reads, number of reads per species proportional to genome length ("24")<br> 3d) "high-coverage, logarithmic" - 240 million reads, number of reads per species proportional to the logarithm of genome length ("24log")</p> <p>4) Training, validation and test sets for the multiclass models. They should be used together with the "pathogenic" read sets hosted at <a href="https://zenodo.org/record/4456857">https://zenodo.org/record/4456857</a>. Here, we share sets for two of the four total classes:<br> 4a) The 'non-pathogen' class is a mixture of "nonpathogenic" biacterial and viral read sets, concatenated and downsampled to the original read number (20M for training, 1.25M for validation and test). The training and validation sets contain mixed-length (25-20bp) simulated subreads (original sets hosted here: <a href="https://zenodo.org/record/4456857">https://zenodo.org/record/4456857</a>). The test set contains 250bp long reads based on the test sets from here: <a href="https://zenodo.org/record/3678563">https://zenodo.org/record/3678563</a> and here: <a href="https://zenodo.org/record/4312525">https://zenodo.org/record/4312525</a>; it was also sorted by species.<br> 4b) Mixed-length versions of the "pathogenic" fungal training and validation sets, prepared by random shortening of the "low-coverage" read sets in the "linear" (_rn_) and "logarithmic" (_rn_*log_) flavours.</p> <p>See also the preprint: <a href="https://www.biorxiv.org/content/10.1101/2021.11.30.470625">https://www.biorxiv.org/content/10.1101/2021.11.30.470625</a></p>
VirHunter: a deep learning-based method for detection of novel RNA viruses in plant sequencing data
<p>This storage contains 2 archives: toy datasets to test the training of the VirHunter and weights of the fully trained VirHunter models for 3 host species (peach, grapevine, sugar beet) and for fragment sizes 500 and 1000. .</p> <p>The toy dataset consists of 3 archived files: 'viruses.fasta', 'host.fasta', 'bacteria.fasta'.</p> <p>'viruses.fasta' contains 10000 randomly selected plant viruses from the virus dataset described in the paper.</p> <p>'host.fasta' consists of peach chromosome 2.</p> <p>'bacteria.fasta' consists of 10 bacterial genomes selected randomly: GCF_000284415, GCF_000590555, GCF_001548055, GCF_002795265, GCF_003330825, GCF_003957805, GCF_005845345, GCF_009176625, GCF_010748935, GCF_014681765</p> <p> </p>
Deciphering the Neurosensory Olfactory Pathway and Associated Neo-Immunometabolic Vulnerabilities Implicated in COVID-Associated Mucormycosis (CAM) and COVID-19 in a Diabetes Backdrop—A Novel Perspective
<p>Raw data files of transcriptomic profiling experiments, which form the basis for our publication (https://www.mdpi.com/2673-4540/3/1/13).</p>
5D-NP-MATER_MDO - Open Dataset for "Novel, High-Resolution, Subtractive Photoresist Formulations for 3D Direct Laser Writing Based on Cyclic Ketene Acetals"
<p>This is the open dataset for the paper: "Marco Carlotti*, Omar Tricinci, Virgilio Mattoli*, Novel, High-Resolution, Subtractive Photoresist Formulations for 3D Direct Laser Writing based on Cyclic Ketene Acetals, Advanced Materials Technologies, On line (2022) [DOI: 10.1002/admt.202101590] "</p> <p>This include the Supplementary Information file ("SI.pdf") , all the source material used for the paper preparation and more. </p> <p>For each folder (sub-dataset) there is a corresponding readme file describing the content and including metadata.</p> <p> </p>
A Novel Dataset of Misinformation Tweets Regarding the CoronaVac Vaccine in Brazil
<p>This dataset was built to analyze the spread of misinformation about CoronaVac in Brazil by using data from Twitter for two specific events: the approval for emergency use in adults over 18 years old (January 17, 2021) and the approval for use in children aged 6 to 17 years (January 20, 2022).</p> <p>We choose to label the original tweets with at least one retweet in the analyzed period. The manual labeling of such tweets was initially performed by two annotators with high knowledge about the dataset and the considered context. In cases in which there was no agreement between the two annotators, a third annotator was considered to define the class of the tweet. </p> <p>The final dataset contains <strong>1,010 tweets from January 17, 2021</strong>, and <strong>816 tweets from January 20, 2022</strong>.</p> <p>This dataset was originally built for a conference paper accepted at BraSNAM 2022. If you make use of the dataset, please also cite the following paper:</p> <p><em>Gabriel P. Oliveira, Beatriz F. Paiva, Ana Paula Couto da Silva, and Mirella M. Moro. Characterizing the Diffusion of Misinformation Regarding the CoronaVac Vaccine in Brazil. In Proceedings of the XI Brazilian Workshop on Social Network Analysis and Mining </em><em>(BraSNAM 2022), 2022.</em></p> <pre><code>@inproceedings{brasnam/OliveiraPSM22, title = {Characterizing the Diffusion of Misinformation Regarding the CoronaVac Vaccine in Brazil}, author = {Gabriel P. Oliveira and Beatriz F. Paiva and Ana Paula Couto da Silva and Mirella M. Moro}, booktitle = {Proceedings of the XI Brazilian Workshop on Social Network Analysis and Mining (BraSNAM)} year = {2022} }</code></pre>
Mapping of samples – Deliverables in WP3: Novel pyrolysis oil from non-food/feed biomass
<p>In the H2020-project BioMates (www.biomates.eu, see chapter 4 “Funding and disclaimer”), RISE produced samples from ablative fast pyrolysis of herbaceous biomass in a TRL 5-plant. A dedicated document provides identifiers for relevant liquid samples and their blends [1]. The document at hand maps it to the substances reported to be used in the public deliverables and public deliverable summaries connected to Work Package 3 “Technology scale up and validation” of the BioMates-project, as far as mapping is not provided in the documents themselves.</p>
Sharing research data and findings relevant to the novel coronavirus (COVID-19) outbreak - Literature sources
<p>The spreadsheet in the present dataset (CSV format) includes the sources considered during the literature review stage for the report: From intent to impact: Investigating the effects of open sharing commitments. Please note that not all sources in this deposit have been referenced in the above-mentioned report and that the report may include additional sources</p>
Sharing research data and findings relevant to the novel coronavirus (COVID-19) outbreak - Survey responses
<p>The spreadsheets in the present dataset (CSV format) include the anonymised responses to our online survey of signatories of the Joint Statement on open research and data sharing. Responses have been split into quantitative responses (i.e., closed survey questions) and qualitative responses (i.e., free text survey questions).</p> <p>This data has been used to inform our final report, which is available in our <a href="https://zenodo.org/communities/data-sharing-in-public-health-emergencies">Zenodo Project Community</a>.</p>
Sharing research data and findings relevant to the novel coronavirus (COVID-19) outbreak - Thematic coding of qualitative research findings
<p>The spreadsheet in the present dataset (CSV format) includes the anonymised thematic coding that has been applied to our interview and literature review findings to inform the preparation of the report: From intent to impact: Investigating the effects of open sharing commitments.</p> <p>The thematic coding has been applied by using <a href="https://www.qsrinternational.com/nvivo-qualitative-data-analysis-software/home">NVivo</a>, a professional qualitative analysis software, and then exported in spreadsheet form for public sharing.</p> <p>Find out more about this project in our dedicated <a href="https://zenodo.org/communities/data-sharing-in-public-health-emergencies">Zenodo project community</a>.</p>
Defining the genes required for survival of Mycobacterium bovis in the bovine host offers novel insights into the genetic basis of survival of pathogenic mycobacteria
<p>Supplementary dataset from "<strong><em>Defining the genes required for survival of Mycobacterium bovis in the bovine host offers novel insights into the genetic basis of survival of pathogenic mycobacteria</em></strong>"</p> <p> </p> <p><strong>Supplementary Figure legends</strong></p> <p><strong>Figure S1. Illustration of the transposon insertions around the <em>M. bovis </em>genome. </strong>Sequencing of the input library showed that transposon insertions were evenly distributed around the genome and 27,419 of the permissible 66,931 thymine–adenine dinucleotide (TA) sites contained an insertion representing an insertion density of ~41%. The outer ring are the genomic coordinates, the blue lines represent transposon insertions and the gray boxes indicate regions of that did not have any insertions. Plot made with Circlize (Gu et al, 2014).</p> <p> </p> <p><strong>Figure S2. Diversity of the output library isolated from lung and thoracic lymph node lesions compared to the input library. </strong>On average, libraries recovered from lung lesions contained 14,456 unique mutants and those recovered from the lymph nodes contained an average of 16,210 unique mutants. Insertion density is represented as a proportion of the TA sites that contained insertions. The numbers on the x-axis refer to the sequencing file from that sample and come from individual animals (Bioproject ID: PRJNA816175, Submission ID: SUB11067380).</p> <p> </p> <p><strong>Figure S3. Volcano plots showing the distribution of log<sub>2</sub> fold-changes and -log<sub>10</sub> of adjusted p-values for representative lung (A) and lymph node (B) samples. </strong>Adjusted p-values (BH-fdr correction) < 0.000001 cluster at the limits of the plot and precision reflects the number of resampling iterations (10,000).</p> <p> </p> <p><strong>Figure S4. Scatterplot of mean log<sub>2</sub> fold change per gene for all lung samples against all thoracic lymph node samples</strong>. Correlation between mean log<sub>2</sub> fold change among genes between the tissues was calculated with Spearman's ranked correlation, = 0.878, p-value < 2.2e-16.</p> <p> </p> <p><strong>Figure S5. Fold-changes caused by transposon insertions in <em>RD1<sup>BCG</sup></em> and <em>RD1<sup>MIC</sup> </em>in the lungs and lymph nodes of infected cattle. </strong>Boxplot for log<sub>2 </sub>fold-changes in genes of the RD1<sup>BCG</sup> region. Samples with adjusted p-values (BH-fdr corrected) <0.05 are indicated with purple points. Gene names highlighted in magenta have fewer than 5 TA sites located in the gene; too few to determine the statistical significance of changes in insertion levels with this method.</p> <p> </p> <p><strong>Supplementary Tables </strong></p> <p><strong>Table S1. Sequencing statistics of the input and output transposon libraries. </strong>The numbers in the column labelled “filename” refers to the sequencing file from that sample and come from individual animals (Bioproject ID: PRJNA816175, Submission ID: SUB11067380).</p> <p> </p> <p><strong>Table S2. Tissues collected and scored for gross pathology. </strong>Tissues from head and neck lymph nodes (from the right and left sub-mandibular lymph nodes, the right and left medial retropharyngeal lymph nodes), thoracic lymph nodes (the right and left bronchial lymph nodes, the cranial tracheobronchial lymph nodes, the cranial and caudal mediastinal lymph nodes) and from lung lesions, were collected and scored.</p> <p> </p> <p><strong>Table S3. Log<sub>2</sub> fold-changes for insertions across the entire genome of <em>M. bovis</em> AF2122/97. </strong>Cells are coloured according to log<sub>2</sub> fold-change. Refer to the text for the gene groups in individual tabs.</p> <p> </p> <p><strong>Table S4. </strong>Custom transposon sequencing primers and adaptors used in sequencing of the transposon libraries.</p> <p> </p>
Summary Statistics from "Meta-GWAS of PCSK9 levels detects two novel loci at APOB and TM6SF2"
<p>GWAMA summary statistics of PCSK9 levels using fixed-effect model. Genome-wide data is given for Europeans with statin adjustment and Europeans without statin treatment only (subset of the population). In addition, locus-wide data of the PCSK9 gene locus for African-Americans without statin treatment is listed.</p> <p>When using this data, please cite: Pott J, Gadin J, Theusch E, et al.. Meta-GWAS of PCSK9 levels detects two novel loci at APOB and TM6SF2. Hum Mol Genet. 2021 Sep 30:ddab279. doi: 10.1093/hmg/ddab279. PMID: 34590679</p> <p>All txt files contain the following columns:</p> <ul> <li>markername</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>ea (effect allele)</li> <li>oa (other allele)</li> <li>eaf (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (number of studies)</li> <li>beta (effect estimate)</li> <li>se (standard error)</li> <li>p (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>phenotype (phenotyp setting)</li> </ul>
Acceleration Research on Novel Photovoltaic Materials
<p>Data and Simulation definiton file for SCAPS1D simulation that are the basis for figures 3-5 of publication DOI:10.1039/d2fd00085g, published in Faraday Discussions (2022)</p> <p>Device structure for the drift-diffusion simulation is: </p> <p>metal back-contact/p-type absorber(1 micron)/n-type buffer layer (30nm)/i-ZnO(80nm)/n-type ZnO(100nm)</p> <p>No interface recombination and no back contact recombination is assumed.</p>
Supplementary Date for a "Novel production of macrocapsules for self-sealing mortar specimens using stereolithographic 3D printers"
<p>This dataset was used for the publication of "<span>Novel production of macrocapsules for self-sealing mortar specimens using stereolithographic 3D printers". </span></p>
Data for: Generation of sanitation system options for urban planning considering novel technologies
<p>This data has been used (1) to quantify the appropriateness of a set of sanitation technologies for a small town (Katarnyia) in Nepal and (2) to generate sanitation system options from the appropriate technologies as an input into strategic sanitation planning using a structured decision making process. For (1), the appropriateness is quantified based on a set of criteria, also called screening criteria. These criteria include technical, socio-demographic, climatic, and institutional aspects and are quantified using uncertainty functions in order to account for the quality and quantity of available input information.</p> <p>The data contains raw data as well as modelling results. The raw data is a compilation of information collected from literature, information collected through a household survey in the small town, field observations. They are all used to describe the screening criteria for the studied sanitation technologies and the small town. Results include: (1) the outcome of the technology appropriateness assessment (technology appropriateness scores); and (2) the sanitation system options (all possible sanitation systems built from the appropriate technologies, and a smaller set of divers and highly appropriate sanitation system options as an input into decision-making).</p>
All-sky information content analysis for novel passive microwave instruments - data
<p>This dataset is the underlying data for the article:</p> <p>Grützun, V., S. A. Buehler, L. Kluft, M. Brath, J. Mendrok, and P. Eriksson (in press, 2018), All-sky Information Content Analysis for Novel Passive Microwave Instruments in the Range from 23.8 GHz up to 874.4 GHz, Atmos. Meas. Tech., doi:10.5194/amt-2017-377. </p> <p>Please refer to that article for a description of the scientific background of the data and to the attached README file for a technical documentation. </p> <p>Contact: Verena Grützun, verena.gruetzun@uni-hamburg.de<br> </p>
Locked Shields Partners Run 23 (LSPR23): A novel IDS dataset from the largest live-fire cybersecurity exercise
<p>IDS Dataset from the Largest Live Fire Cybersecurity Exercise Using Virtual Blue Team Network Traffic.<br><br></p> <ul> <li> <p>LSPR23 is derived from Locked Shields 2023, a major live-fire cyber defense exercise.</p> </li> <li> <p>LSPR23 includes ~16M network flows, of which ~1.6M are labeled malicious.</p> </li> </ul> <p> </p> <p>Please cite our research article:"LSPR23: A novel IDS dataset from the largest live-fire cybersecurity exercise" when using our dataset:<br>https://doi.org/10.1016/j.jisa.2024.103847<br><br><br></p>
Mapping a novel metric for Flash Flood Recovery using Interpretable Machine Learning
<p>This data is supplementary to the paper titled "Mapping a novel metric for Flash Flood Recovery using Interpretable Machine Learning". The file contains the main results.<br><br>For any queries, please visit <a href="https://hydrosense.iitd.ac.in" target="_blank" rel="noopener">Hydrosense Lab (IIT Delhi)</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.