Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
468
datasets available to search
ShareScore release 0.9.0
Dataset results
468 results for “MERS”
Quantum mechanical electronic and geometric parameters for DNA k-mers as features for machine learning
<p>With the development of advanced predictive modelling techniques, we are witnessing a steep increase in model development initiatives in genomics that employ high-end machine learning methodologies. Of particular interest are models that predict certain genomic or biological characteristics based solely on DNA sequence information. These models, however, treat the DNA sequence as a mere collection of four, A, T, G and C, letters, thus dismissing the past physico-chemical advancements in science that can enable the use of more intricate information about nucleic acid sequences. Here, we provide a comprehensive database of quantum mechanical and geometric features for all the permutations of 7-meric DNA in their representative B, A and Z conformations. The database is generated by employing the applicable high-cost and time-consuming quantum mechanical methodologies. This can thus make it seamless to associate a wealth of novel molecular features to any DNA sequence, by scanning it with a matching k-meric window and pulling the pre-computed values from our database for further use in modelling. We demonstrate the usefulness of our deposited features through their exclusive use in developing a model for A to C mutation rate constants.</p> <p>The DNA k-mer quantum mechanical parameters can also be found <a href="https://github.com/SahakyanLab/DNAkmerQM" target="_blank" rel="noopener">https://github.com/SahakyanLab/DNAkmerQM</a>, the corresponding research and development code from <a href="https://github.com/SahakyanLab/NucleicAcidsQM" target="_blank" rel="noopener">https://github.com/SahakyanLab/NucleicAcidsQM</a>, and the associated pre-print from <a href="https://doi.org/10.1101/2023.01.25.525597" target="_blank" rel="noopener">https://doi.org/10.1101/2023.01.25.525597</a>.</p>
N2 fixation rates at Mer Bleue bog, Ontario, Canada
<p>This dataset contains bi-weekly measurements of N2 fixation rates and ancillary measurements at the Mer Bleue bog in Ontario, Canada. A detailed description of the dataset can be found in the accompanying publication Zivkovic et al. (Journal of Geophysical Research - Biogeosciences, in review).</p>
Taxonomic classification based on k-mers
<p>DNA sequencing provides the possibility to obtain complete genomic DNA from environmental samples without the need for laboratory microbiological cultures. To this end, metagenomics, the direct DNA sequencing from microbial communities, has changed radically the field of microbiology, by unearthing a broad space of the planet’s microbial diversity, much of which remains unknown. Metagenomic approaches have become standard methods for identifying the biodiversity and the gene or metabolomic functionalities of bacterial and archaeal communities, with many applications not only in microbial ecology but also in public health as in clinical diagnostics and detection of pathogens. Yet, the decrease in the cost of high throughput sequencing and the great amount of microbial data produced every day, highlight one of the main biological questions: the taxonomic classification of metagenomic short reads. Which organisms are contained in a sample? Are there any features that can be used to identify them?</p> <p>To this end, many algorithms have been developed that achieve high speed, by counting k-mers, short sequence substrings of fixed-length k. In this way for the provided input sequences, a list of features can be computed that describes each one of them. Subsequently, the question now reforms to how can the produced k-mers be used for the taxonomic classification of the input sequences.</p> <p>Sample processing, sequencing, and core amplicon data analysis were performed by the Earth Microbiome Project (www.earthmicrobiome.org), and all amplicon sequence data and metadata have been made public through the EMP data portal (qiita.microbio.me/emp):</p> <ul> <li>Thompson, L. R., Sanders, J. G., McDonald, D., Amir, A., …, Jansson, J. K., Gilbert, J. A., Knight, R., & The Earth Microbiome Project Consortium. (2017). A communal catalogue reveals Earth’s multiscale microbial diversity. Nature, 551:457-463. doi:10.1038/nature24621.</li> </ul>
Photos d'un concours de pêche sur la plage de Boulogne-sur-Mer
<p>Photos prises lors d'un concours de Surf-Casting le 7 mai sur la plage de Boulogne-sur-Mer. Le premier plan des photos montre la plage, sur laquelle sont installés des pêcheurs avec leurs cannes à pêche. L'arrière plan montre la mer ou le port de la ville en fonction de l'angle de prise de vue. </p>
Photos de la digue Carnot à Boulogne-sur-Mer
<p>Ces trois photos ont été prises au mois de mai 2017 sur la digue Carnot à Boulogne-sur-Mer. </p> <p>Elles montrent les traces des activités humaines : restes de coquillage destinés à servir d'appât, arrêté préfectoral arraché, poubelle remplie de détritus. </p>
30-mer mappable regions in the human hg19 genome
<p>Knowing where reads can uniquely map in the genome is useful for nascent RNA assays, both in statistical calculations and to make predictions.</p> <p>The dataset was created using the bowtie 1 aligner. The genome was windows at 30 basepair genomic intervals and mapped back to the genome. If the read maps to more than one place, the read is thrown away. Therefore the regions captured in the dataset are regions that any read at least 30 basepairs long will map to uniquely. The shell script originally used to create this dataset has been lost.</p>
Multi-reference genome and K-mer based association mapping in Zymoseptoria tritici
<p>Data tables for a study of multi-reference genome and K-mer based association mapping of the fungal wheat pathogen <em>Zymoseptoria tritici</em></p>
Microelectrode register (MER) data from Deep Brain Stimulation (DBS) surgery in Parkinson's disease patients
<p>MER data consist of brain signal in different depths when DBS surgery is being done. In each depth, a data file is created, with different duration depending on the depth, and up to three channels.</p> <p>Data come from 14 patients (9 males and 5 females), they are anonymized and labelled from P1 to P14. They correspond to patients in age 65.1 +- 5.6 years.</p> <p>Data are organized in STN-IN and STN-OUT (different depths in each folder), subthalamus-in, and subthalamus-out since the STN area is the target area when implanting a DBS. Classification in STN-IN and STN-OUT was made by the neurophysiologists and surgeons.</p> <p>Data were recorded for left and right lobes, 8 patients in left and right lobe, 1 patient in right lobe, and 5 patients in left lobe.</p> <p>The format is mat file (MATLAB file)</p> <p>Data sampling frequency is 12kHz.</p> <p>No filtering or data processing was made, they are directly obtained from the MER acquisition system.</p> <p> </p>
Plasmer database for k-mer and genomic features
<p>This is the inital version v1.0 of Plasmer database for k-mer and genomic features.</p> <p>Download and extract the package, and provide the absolute path to the Plasmer command line.</p> <p> </p> <p>For more information about Plasmer, please refer to our GitHub repository at: <a href="https://github.com/nekokoe/plasmer">https://github.com/nekokoe/plasmer</a></p>
Weighted k-mer datasets
<p>These are the datasets used in the experiments of the paper: <em>On Weighted k-mer Dictionaries</em> - Giulio Ermanno Pibiri. Algorithms for Molecular Biology, <strong>18</strong>, Article number: 3 (2023) DOI: <a href="https://almob.biomedcentral.com/articles/10.1186/s13015-023-00226-2">10.1186/s13015-023-00226-2</a>. (A preliminary version of the paper has been published in WABI 2022: <a href="https://doi.org/10.4230/LIPIcs.WABI.2022.9">10.4230/LIPIcs.WABI.2022.9</a>.)</p>
Data from: A k-mer-based approach for phylogenetic classification of taxa in environmental genomic data
<p>In the age of genome sequencing, whole genome data is readily and frequently generated, leading to a wealth of new information that can be used to advance various fields of research. New approaches, such as alignment-free phylogenetic methods that utilize <em>k</em>-mer-based distance scoring, are becoming increasingly popular given their ability to rapidly generate phylogenetic information from whole genome data. However, these methods have not yet been tested using environmental data, which often tends to be highly fragmented and incomplete. Here we compare the results of one alignment-free approach (which utilizes the D<sup>2</sup> statistic) to traditional multi-gene maximum likelihood trees in three algal groups that have high-quality genome data available. In addition, we simulate lower-quality, fragmented genome data using these algae to test method robustness to genome quality and completeness. Finally, we apply the alignment-free approach to environmental metagenome assembled genome data of unclassified Saccharibacteria and Trebouxiophyte algae, and single-cell amplified data from uncultured marine stramenopiles to demonstrate its utility with real datasets. We find that in all instances, the alignment-free method produces phylogenies that are comparable, and often more informative, than those created using the traditional multi-gene approach. The k-mer-based method performs well even when there is significant missing data, that includes marker genes traditionally used for tree reconstruction. Our results demonstrate the value of alignment-free approaches for classifying novel, often cryptic or rare, species, that may not be culturable or are difficult to access using single-cell methods but fill important gaps in the tree of life.</p>
Genetic networks datasets: ten books by Gustave Roud, "En mer" by Bernard Comment
<p>In genetic criticism, scholarly editing, authorial philology and, more generally, for the study of authorial manuscripts and writing processes it is essential to order and classify the textual witnesses and their relationships. The two datasets documents so-called ‘genetic networks’, representations of the genetic entities (witnesses, publications, dossiers) and of their relationships, modelled according to the GENO 1.0 ontology. The datasets contain genetic networks of the works of two Swiss authors: the main publications of Gustave Roud (1897-1976) and the short story “En mer” by Bernard Comment (1960).</p>
Safety, Tolerability and Immunogenicity of INO-4700 for MERS-CoV in Healthy Volunteers
ClinicalTrials.gov study NCT04588428. IPD Sharing: YES. Countries: 3. Publications: 1.
A Study to Determine the Efficacy, Safety and Tolerability of Aztreonam-Avibactam (ATM-AVI) ± Metronidazole (MTZ) Versus Meropenem (MER) ± Colistin (COL) for the Treatment of Serious Infections Due to
ClinicalTrials.gov study NCT03329092. IPD Sharing: YES. Countries: 21. Publications: 4.
Safety, Tolerability and Immunogenicity of Vaccine Candidate MVA-MERS-S
ClinicalTrials.gov study NCT03615911. IPD Sharing: NO. Countries: 1. Publications: 10.
Data from: A k-mer-based approach for phylogenetic classification of taxa in environmental genomic data
Open the record for dataset details and reuse information.
Sars-Cov-2 and Mers sequences from human host with no unknown characters
Open the record for dataset details and reuse information.
History of coronavirus naming during the three zoonotic outbreaks in relation to virus taxonomy and diseases caused by these viruses. According to the current international classification of diseases49, MERS and SARS are classified as 1D64 and 1D65, respectively. in The species Severe acute respiratory syndromerelated coronavirus: classifying 2019-nCoV and naming it SARS-CoV-2
History of coronavirus naming during the three zoonotic outbreaks in relation to virus taxonomy and diseases caused by these viruses. According to the current international classification of diseases49, MERS and SARS are classified as 1D64 and 1D65, respectively.
BAFF 60-mer and BAFF 60-mer-dissociating activities in serum, cord blood and cerebrospinal fluid
<p>This dataset is related to "BAFF 60-mer, and differential BAFF 60-mer dissociating activities in human serum, cord blood and cerebrospinal fluid" (Eslami M, Meinl E, Eibel H, Willen L, Donzé O, Distl O, Schneider H, Speiser DE, Tsiantoulas D, Yalkinoglu Ö, Samy E, Schneider P).</p>
MER Opportunity and Spirit Rovers Pancam Images Labeled Data Set
<p><strong>Introduction</strong></p> <p>The data set is based on 3,004 images collected by the Pancam instruments mounted on the Opportunity and Spirit rovers from NASA's Mars Exploration Rovers (MER) mission. We used rotation, skewing, and shearing augmentation methods to increase the total collection to 70,864 (see Image Augmentation section for more information). Based on the <a href="https://merdatacatalog.com/survey">MER Data Catalog User Survey</a> [1], we identified 25 classes of both scientific (e.g. soil trench, float rocks, etc.) and engineering (e.g. rover deck, Pancam calibration target, etc.) interests (see Classes section for more information). The 3,004 images were labeled on <a href="https://www.zooniverse.org/">Zooniverse platform</a>, and each image is allowed to be assigned with multiple labels. The images are either 512 x 512 or 1024 x 1024 pixels in size (see Image Sampling section for more information).</p> <p><strong>Classes</strong></p> <p>There is a total of 25 classes for this data set. See the list below for class names, counts, and percentages (the percentages are computed as count divided by 3,004). Note that the total counts don't sum up to 3,004 and the percentages don't sum up to 1.0 because each image may be assigned with more than one class. </p> <ul> <li>Class name, count, percentage of dataset</li> <li>Rover Deck, 222, 7.39%</li> <li>Pancam Calibration Target, 14, 0.47%</li> <li>Arm Hardware, 4, 0.13%</li> <li>Other Hardware, 116, 3.86%</li> <li>Rover Tracks, 301, 10.02%</li> <li>Soil Trench, 34, 1.13%</li> <li>RAT Brushed Target, 17, 0.57%</li> <li>RAT Hole, 30, 1.00%</li> <li>Rock Outcrop, 1915, 63.75%</li> <li>Float Rocks, 860, 28.63%</li> <li>Clasts, 1676, 55.79%</li> <li>Rocks (misc), 249, 8.29%</li> <li>Bright Soil, 122, 4.06%</li> <li>Dunes/Ripples, 1000, 33.29%</li> <li>Rock (Linear Features), 943, 31.39%</li> <li>Rock (Round Features), 219, 7.29%</li> <li>Soil, 2891, 96.24%</li> <li>Astronomy, 12, 0.40%</li> <li>Spherules, 868, 28.89%</li> <li>Distant Vista, 903, 30.23%</li> <li>Sky, 954, 31.76%</li> <li>Close-up Rock, 23, 0.77%</li> <li>Nearby Surface, 2006, 66.78%</li> <li>Rover Parts, 301, 10.02%</li> <li>Artifacts, 28, 0.93%</li> </ul> <p><strong>Image Sampling</strong></p> <p>Images in the MER rover Pancam archive are of sizes ranging from 64x64 to 1024x1024 pixels. The largest size, 1024x1024, was by far the most common size in the archive. For the deep learning dataset, we elected to sample only 1024x1024 and 512x512 images as the higher resolution would be beneficial to feature extraction.</p> <p>In order to ensure that the data set is representative of the total image archive of 4.3 million images, we elected to sample via "site code". Each Pancam image has a corresponding two-digit alphanumeric "site code" which is used to track location throughout its mission. Since each "site code" corresponds to a different general location, sampling a fixed proportion of images taken from each site ensure that the data set contained some images from each location. In this way, we could ensure that a model performing well on this dataset would generalize well to the unlabeled archive data as a whole. We randomly sampled 20% of the images at each site within the subset of Pancam data fitting all other image criteria, applying a floor function to non-whole number sample sizes, resulting in a dataset of 3,004 images.</p> <p><strong>Train/validation/test sets split</strong></p> <p>The 3,004 images were split into train, validation, and test data sets. The split was done so that roughly 60, 15, and 25 percent of the 3,004 images would end up as train, validation, and test data sets respectively, while ensuing that images from a given site are not split between train/validaiton/test data sets. This resulted in 1,806 train images, 456 validation images, and 742 test images. </p> <p><strong>Augmentation</strong></p> <p>To augment the images in train and validation data sets (note that images in the test data set were not augmented), three augmentation methods were chosen that best represent transformations that could be realistically seen in Pancam images. The three augmentations methods are rotation, skew, and shear. The augmentation methods were applied with random magnitude, followed by a random horizontal flipping, to create 30 augmented images for each image. Since each transformation is followed by a square crop in order to keep input shape consistent, we had to constrict the magnitude limits of each augmentation to avoid cropping out important features at the edges of input images. Thus, rotations were limited to 15 degrees in either direction, the 3-dimensional skew was limited to 45 degrees in any direction, and shearing was limited to 10 degrees in either direction. Note that augmentation was done only on training and validation images. </p> <p><strong>Directory Contents</strong></p> <ul> <li>images: contains all 70,864 images</li> <li>train-set-v1.1.0.txt: label file for the training data set</li> <li>val-set-v1.1.0.txt: label file for the validation data set</li> <li>test-set-v1.1.0.txt: label file for the testing data set</li> </ul> <p>Images with relatively short file names (e.g., 1p128287181mrd0000p2303l2m1.img.jpg) are original images, and images with long file names (e.g., 1p128287181mrd0000p2303l2m1.img.jpg_04140167-5781-49bd-a913-6d4d0a61dab1.jpg) are augmented images. The label files are formatted as "Image name, Class1, Class2, ..., ClassN".</p> <p> </p> <p><strong>Reference</strong></p> <p>[1] S.B. Cole, J.C. Aubele, B.A. Cohen, S.M. Milkovich, and S.A. Shields, Identifying Community Needs for a Mars Exploration Rovers (MER), Daata Catalog, 51st Lunar and Planetary Science Conference (LPSC), 2020.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.