Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
468
datasets available to search
ShareScore release 0.9.0
Dataset results
468 results for “MERS”
Data from: Association between severity of MERS-CoV infection and incubation period
We analyzed data for 170 patients in South Korea who had laboratory-confirmed infection with Middle East respiratory syndrome coronavirus. A longer incubation period was associated with a reduction in the risk for death (adjusted odds ratio/1-day increase in incubation period 0.83, 95% credibility interval 0.68–1.03).
Data from: Tracing coco de mer's reproductive history: pollen and nutrient limitation reduce fecundity
Habitat degradation can reduce or even prevent the reproduction of previously abundant plant species. To develop appropriate management strategies, we need to understand the reasons for reduced recruitment in degraded ecosystems. The dioecious coco de mer palm (Lodoicea maldivica) produces by far the largest seeds of any plant. It is a keystone species in an ancient palm forest that occurs only on two small islands in the Seychelles, yet contemporary rates of seed production are low, especially in fragmented populations. We developed a method to infer the recent reproductive history of female trees from morphological evidence present on their inflorescences. We then investigated the effects of habitat disturbance and soil nutrient conditions on flower and fruit production. The 57 female trees in our sample showed a 19.5-fold variation in flower production among individuals over a seven-year period. Only 77.2% of trees bore developing fruits (or had recently shed fruits), with the number per tree ranging from zero to 43. Flower production was positively correlated with concentrations of available soil nitrogen and potassium, and did not differ significantly between closed and degraded habitat. Fruiting success was positively correlated with pollen availability, as measured by numbers and distance of neighbouring male trees. Fruit-set was lower in degraded habitat than in closed forest, while the proportion of abnormal fruits that failed to develop was higher in degraded habitat. Seed size recorded for a large sample of seeds varied widely, with fresh weights ranging from 1 to 18 kg. Shortages of both nutrients and pollen appear to limit seed production of Lodoicea in its natural habitat, with these factors affecting different stages of the reproductive process. Flower production varies widely amongst trees, while seed production is especially low in degraded habitat. The size of seeds is also very variable. We discuss the implications of these findings for managing this ecologically and economically important species.
Leveraging basecaller's move table to generate a lightweight k-mer model
<p>The ONT RNA004 dataset that is used to create a 5-mer model using the basecaller's movetable.</p> <p>The dataset is a sub-sample extracted from a dataset that sequenced Universal Human Reference RNA.</p> <p>The specification of the bio sample is here</p> <p>https://www.agilent.com/cs/library/usermanuals/public/740000.pdf</p>
Genomic datasets used for evaluation of k-mer representations and indexes
<p>This record contains genomic datasets, including subsampled <em>k</em>-mer sets for some datasets (files with names containing _subsampled_). Namely, it provides the following datasets:</p> <ul> <li>Two <em>E. coli</em> pan-genomes, obtained as the union of the <em>E. coli</em> genomes from the 661k collection. One contains <em>all</em> genomes (without quality filtering) and for the other (<em>HQ</em>) we applied high-quality filtering.</li> <li><em>S. pneumoniae</em> pan-genome: 616 genomes, as provided in RASE DB <em>S. pneumoniae</em> <a href="https://github.com/c2-d2/rase-db-spneumoniae-sparc/">https://github.com/c2-d2/rase-db-spneumoniae-sparc/</a></li> <li><em>SARS-CoV-2</em> pan-genome, downloaded from GISAID <a href="https://gisaid.org/">https://gisaid.org/</a> (access upon registration) on Jan 25, 2023 (GISAID version 2023/01/23, 14,682,066 genomes, 430 Gbp).</li> <li>Metagenomic sample SRS063932 (Illumina raw reads) of human microbiome with accession SRX023459, download from <a href="https://www.hmpdacc.org/hmp/HMASM/">https://www.hmpdacc.org/hmp/HMASM/</a>. The fastq files were converted to FASTA files using `seqtk seq -A -C`.</li> <li>Human RNA-seq Illumina raw reads with accession SRX348811, downloaded using the prefetch tool from the SRA toolkit and then converted into the FASTA format by<br>`fastq-dump --split-3 --fasta`.</li> <li>Human genome Illumina raw reads with accession SRX016231, downloaded using the prefetch tool from the SRA toolkit and then converted into the FASTA format by<br>`fastq-dump --split-3 --fasta`.</li> <li>Human genome assembly chm13v2.0 (T2T), downloaded from <a href="https://s3-us-west-2.amazonaws.com/human-pangenomics/T2T/CHM13/assemblies/analysis_set/chm13v2.0.fa.gz">https://s3-us-west-2.amazonaws.com/human-pangenomics/T2T/CHM13/assemblies/analysis_set/chm13v2.0.fa.gz</a>.</li> <li>Two MiniKraken datasets (4GB and 8GB), downloaded from <a href="https://ccb.jhu.edu/software/kraken/">https://ccb.jhu.edu/software/kraken/</a>, with the 31-mers dumped using Jellyfish 1.1.12.</li> </ul> <p>The resulting FASTA files (apart from the human genome assembly chm13v2.0 and MiniKraken datasets) were converted to unitigs by GGCAT v1.1.0<br>by `ggcat build -k {kmer-size} -m 200 -j 5 -s {min-freq} -o {preprocessed_unitigs} {input_FASTA}`, where we used $k=128$ and `{min-freq}`=1 for pan-genomes and $k=32$ and `{min-freq}`=2 for dataset from raw reads.</p> <p>Finally, the subsampled files `{dataset}_subsampled_k{$k$}_r0.1.fa.xz` contain 10% randomly chosen distinct canonical $k$-mers from the whole $k$-mer set of the given dataset. The FASTA file contains one subsampled <em>k</em>-mer per sequence.</p>
15-mer rank table for use with Fuzzion2
<p>The Fuzzion2 program requires a k-mer rank table. This is a 4-GB binary file holding the frequency ranks of 15-mers in the GRCh38 human reference genome. See <a href="https://github.com/stjude/fuzzion2">https://github.com/stjude/fuzzion2</a> for details.</p>
Dataset for journal article: "Acid-base-induced fac -> mer isomerization of luminescent iridium(III) complexes"
<p>Anastasia Yu. Gitlina, Farzaneh Fadaei-Tirani, Albert Ruggi, Carolina Plaice, Kay Severin*<br> Acid-base-induced <em>fac→mer</em> isomerization of luminescent iridium(III) complexes<br> <em>Chem. Sci.</em> <strong>2022</strong>, DOI: 10.1039/d2sc02808e</p> <p>The dataset contains the following raw data - NMR, HRMS, XRD, CD, HPLC, photophysical characterization (emission, excitation, absorption, emission lifetimes, emission quantum yields). A short video showing the isomerization from fac- to mer-Ir(ppy)3 in the well plate is also included to the data folder.</p>
MER dataset im2latexv2 - Part 1
<h1>Mathematical Expression Recognition Dataset im2latexv2 - Part 1</h1> <p>This repository contains Part 1 of the im2latexv2 dataset presented in the paper <em>MathNet: A Data-Centric Approach for Printed Mathematical Expression Recognition.</em></p> <p>The dataset is an enhanced version of the im2latex-100k dataset. It uses a novel LaTeX normalization process and 61 rendering environments to make the dataset more realistic.</p> <p>Please also download <a href="../records/11296280">Part 2</a> of the im2latexv2 dataset (doi: 10.5281/zenodo.11296280) and copy the subfolders in the folder of Part 1.</p> <p>To unpack all images, please use the unpack_im2latexv2.py script.<br><br>The CSV files have the following structure:</p> <table> <tbody> <tr> <td>formula</td> <td>images</td> <td> </td> <td> </td> </tr> <tr> <td>tokenized formula (tokens separated by white spaces)</td> <td>path to image with rendering env 1</td> <td>path to image with rendering env 2</td> <td>....</td> </tr> </tbody> </table>
MER dataset im2latexv2 - Part 2
<h1>Mathematical Expression Recognition Dataset im2latexv2 - Part 2</h1> <p>This repository contains Part 1 of the im2latexv2 dataset presented in the paper <em>MathNet: A Data-Centric Approach for Printed Mathematical Expression Recognition.</em></p> <p>The dataset is an enhanced version of the im2latex-100k dataset. It uses a novel LaTeX normalization process and 61 rendering environments to make the dataset more realistic.</p> <p>You also need to download <a href="../records/11230382">Part 1</a> (doi: 10.5281/zenodo.11230382) to use the im2latexv2 dataset.</p>
MER dataset realFormula
<h1>Mathematical Expression Recognition Dataset realFormula</h1> <p>This repository contains the realFormula dataset presented in the paper <em>MathNet: A Data-Centric Approach for Printed Mathematical Expression Recognition.</em></p> <p>The dataset contains 121 manual annotated mathematical expressions from arXiv papers.</p> <p>The CSV files have the following structure:</p> <table> <tbody> <tr> <td>formula</td> <td>images</td> </tr> <tr> <td>tokenized formula (tokens separated by white spaces)</td> <td>image name</td> </tr> </tbody> </table>
Data and code for, "Predicting self-assembly of sequence-controlled copoly- mers with stochastic sequence variation"
Open the record for dataset details and reuse information.
turketwh/Prx_3-merSVM_data: First release of Prx 3-mer SVM data set
<p>First release of Prx 3-mer SVM data set to accompany manuscript submissions.</p>
General and Human Specific (sK2-sK6) Universal k-mer Sets
<p>Data sets produced for the publication<strong> Practical universal k-mer sets for minimizer schemes</strong><strong>. </strong>Dan DeBlasio, Fiyinfoluwa Gbosibo, Carl Kingsford , and Guillaume Marçais presented at ACM-BCB 2019. Intended to be used with the ftrie library at https://github.com/Kingsford-Group/remuval</p>
Human Specific (sK7-sK10) Universal k-mer Sets
<p>Data sets produced for the publication<strong>Practical universal k-mer sets for minimizer schemes</strong><strong>. </strong>Dan DeBlasio, Fiyinfoluwa Gbosibo, Carl Kingsford , and Guillaume Marçais presented at ACM-BCB 2019. Intended to be used with the ftrie library at https://github.com/Kingsford-Group/remuval</p>
Innate and adaptive immune genes associated with MERS-CoV infection in dromedaries
<p>The recent SARS-CoV-2 pandemic has refocused attention to the betacoronaviruses, only eight years after the emergence of another zoonotic betacoronavirus, the Middle East respiratory syndrome coronavirus (MERS-CoV). While the wild source of SARS-CoV-2 may be disputed, for MERS-CoV, dromedaries are considered as source of zoonotic human infections. Testing 100 immune- response genes in 121 dromedaries from United Arab Emirates (UAE) for potential association with present MERS-CoV infection, we identified candidate genes with important functions in the adaptive, MHC-class I (HLA-A-24-like) and II (HLA-DPB1-like), and innate immune response (PTPN4, MAGOHB), and in cilia coating the respiratory tract (DNAH7). </p>
Phylo-k-mers databases for SHERPAS
<p>SHERPAS is a new program to identify novel recombinant sequences in a large collection of viral sequences, and to provide a first estimate of their recombinant structure. SHERPAS is much faster than other softwares for recombination detection; its main feature is the use of a pre-computed database of "phylogenetically-informed k-mers" (or phylo-k-mers). The computation of this phylo-k-mer database is a heavy computational step, but it only needs to be executed once for a given reference alignment.</p> <p>A phylo-k-mer database can be built from any reference alignment, and a phylogenetic tree built from that alignment, using RAPPAS2 (<a href="https://github.com/phylo42/rappas2">https://github.com/phylo42/rappas2</a>). We propose here three ready-to-use databases, for three reference alignments:<br> -An alignment of 167 sequences of the pol region of the HIV genome, provided with the program SCUEAL, accessible at <a href="https://github.com/spond/SCUEAL/blob/master/data/pol2009.nex">https://github.com/spond/SCUEAL/blob/master/data/pol2009.nex</a><br> -An alignment of 339 sequence of the whole HBV genome, provided with the programm jpHMM, accessible at <a href="http://jphmm.gobics.de/download.html">http://jphmm.gobics.de/download.html</a>.<br> -An alignment of 881 sequences of the whole HIV genome, also provided with jpHMM, accessible at <a href="http://jphmm.gobics.de/download.html">http://jphmm.gobics.de/download.html</a>.</p> <p>For each of these alignments, we provide a .zip file containing three files: The phylo-k-mer database (.rps file), the reference phylogenetic tree used to build the database (.tree file), and a table associating each reference sequence to a strain of the virus (.csv file). The details of the construction of the database, the construction of the tree, as well as the origin of the information reported in the table, can be found in the Supplementary Materials associated with the original Bioinformatics publication.</p>
KTU: K-mer Taxonomic Units improve the biological relevance of amplicon sequence variant microbiota data
<p>Testing datasets and files for the KTU algorithm</p>
K-mer collision statistics (BLEND: A Fast, Memory-Efficient, and Accurate Mechanism to Find Fuzzy Seed Matches in Genome Analysis)
<p>This dataset contains 1,077 FASTA files and CSV files. Each FASTA file includes 25-character long sequences similar to each other.</p> <p>We have a CSV file for each tool (i.e., minimap2 and BLEND) and configuration (i.e., different number of neighbors in BLEND). CSV files include the non-identical k-mer pairs (16-mers) that generate the same hash value (i.e., collisions). These k-mers are extracted from sequences that are similar to each other. In each line, we show the hash value of the k-mers, the actual sequene pairs that the k-mers are extracted from, k-mer pairs that generate the same hash value, and the edit distance between these k-mers.</p> <p> </p>
Clinical Study to Evaluate the Safety and Effectiveness of MER® Stents in Carotid Revascularisation.
ClinicalTrials.gov study NCT03133429. IPD Sharing: NO. Countries: 1. Publications: 6.
A Clinical Trial to Determine the Safety and Immunogenicity of Healthy Candidate MERS-CoV Vaccine (MERS002)
ClinicalTrials.gov study NCT04170829. IPD Sharing: NO. Countries: 1. Publications: 1.
MERS-CoV Infection tReated With A Combination of Lopinavir /Ritonavir and Interferon Beta-1b
ClinicalTrials.gov study NCT02845843. IPD Sharing: UNDECIDED. Countries: 1. Publications: 4.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.