Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
608
datasets available to search
ShareScore release 0.9.0
Dataset results
608 results for “ensembles”
Long-term trends of ambient nitrate (NO3-) concentrations across China based on ensemble machine-learning models
<p>The monthly NO3- concentrations across China during 2005-2015</p>
Conformational Ensembles of Non-Coding Elements in the SARS-CoV-2 Genome from Molecular Dynamics Simulations
<p>MD trajectories of SL1, SL2, SL2+SL3, SL4, and SL5a elements in SARS-CoV-2 5'UTR</p>
LyNoS: Mediastinal lymph nodes segmentation using 3D convolutional neural network ensembles and anatomical prior guiding
Open the record for dataset details and reuse information.
Extraction of periodic signals in GNSS vertical coordinate time series using adaptive Ensemble Empirical Modal Decomposition method
Open the record for dataset details and reuse information.
Extreme events changes over China under 1.5-4°C global warming targets: projected by an ensemble of regional climate model simulations
<p>This file is for the upload of data for 2019JD031057R.</p>
Dataset for "EA-ERT: a new ensemble approach to convert ERT data to soil moisture"
<p>This dataset supports the research study "EA-ERT: a new ensemble approach to convert ERT data into soil moisture" by B. Loiseau, S. D. Carrière, N. K. Martin-StPaul, R. Clément, C. Champollion, V. Mercier, J. Thiesson, S. Pasquet, C. Doussan, T. Hermans and D. Jougnot.</p> <p>We provide electrical resistivity tomograph (ERT) time-lapse data for the two sites studied (Avignon and Larzac), with associated volumetric water contents measured by the probes in the field. Meanwhile, we also offer the R code for the calculations of each step of the EA-ERT method. </p>
Cocaine-Induced Gene Regulation in D1 and D2 Neuronal Ensembles of the Nucleus Accumbens Revealed by Single-Cell RNA Sequencing
GEO Series GSE297372. Mus musculus. 16 samples. Type: Expression profiling by high throughput sequencing.
RNA sequencing from neural ensembles activated during fear conditioning in the mouse temporal association cortex
GEO Series GSE85128. Mus musculus. 8 samples. Type: Expression profiling by high throughput sequencing.
MRLxSM eQTL in Liver by RMA on Ensembl transcripts
GEO Series GSE25322. Mus musculus. 300 samples. Type: Expression profiling by array.
Ensemble Approach to Building Mercer Kernels
This paper presents a new methodology for automatic knowledge driven data mining based on the theory of Mercer Kernels, which are highly nonlinear symmetric positive definite mappings from the original image space to a very high, possibly infinite dimensional feature space. We describe a new method called Mixture Density Mercer Kernels (MDMK) to learn kernel function directly from data, rather than using pre-defined kernels. These data adaptive kernels can encode prior knowledge in the kernel using a Bayesian formulation, thus allowing for physical information to be encoded in the model. Specifically, we demonstrate the use of the algorithm in situations with extremely small samples of data. We compare the results with existing algorithms on data from the Sloan Digital Sky Survey (SDSS) and demonstrate the method’s superior performance against standard methods. The results show that the Mixture Density Mercer Kernel described here outperforms tree-based classification in distinguishing high-redshift galaxies from low redshift galaxies by approximately 16% on test data, bagged trees by approximately 7%, and bagged trees built on a much larger sample of data by approximately 2%. The code for these experiments has been generated with the AutoBayes tool, which automatically generates efficient and documented C/C++ code from abstract statistical model specifications. The core of the system is a schema library which contains templates for learning and knowledge discovery algorithms like different versions of EM, or numeric optimization methods like conjugate gradient methods. The template instantiation is supported by symbolic algebraic computations, which allows AutoBayes to find closed form solutions and, where possible, to integrate them into the code.
Ensemble Data Mining Methods
Ensemble Data Mining Methods, also known as Committee Methods or Model Combiners, are machine learning methods that leverage the power of multiple models to achieve better prediction accuracy than any of the individual models could on their own. The basic goal when designing an ensemble is the same as when establishing a committee of people: each member of the committee should be as competent as possible, but the members should be complementary to one another. If the members are not complementary, i.e., if they always agree, then the committee is unnecessary---any one member is sufficient. If the members are complementary, then when one or a few members make an error, the probability is high that the remaining members can correct this error. Research in ensemble methods has largely revolved around designing ensembles consisting of competent yet complementary models.
Key Real-World Applications of Classifier Ensembles
Broad classes of statistical classification algorithms have beendeveloped and applied successfully to a wide range of real worlddomains. In general, ensuring that the particular classificationalgorithm matches the properties of the data is crucial inproviding results that meet the needs of the particular applicationdomain. One way in which the impact of this algorithm/applicationmatch can be alleviated is by using ensembles of classifiers, wherea variety of classifiers (either different types of classifiers ordifferent instantiations of the same classifier) are pooled before afinal classification decision is made. Intuitively, classifierensembles allow the different needs of a difficult problem to behandled by classifiers suited to those particular needs.Mathematically, classifier ensembles provide an extra degree offreedom in the classical bias/variance tradeoff, allowing solutionsthat would be difficult (if not impossible) to reach with only asingle classifier. Because of these advantages, classifier ensembles have been applied to many difficult real world problems. In this paper, we surveyselect applications of ensemble methods to problems that havehistorically been most representative of the difficulties inclassification. In particular, we survey applications of ensemblemethods to remote sensing, person recognition, one vs. allrecognition, and medicine.
Ensemble X-Ray Variability of AGN in 2XMMi-DR3
The X-ray variability of active galactic nuclei (AGN) has been most often investigated with studies of individual, nearby sources, and only a few ensemble analyses have been applied to large samples in wide ranges of luminosity and redshift. In their study, the authors aimed to determine the ensemble variability properties of two serendipitously selected AGN samples extracted from the catalogs of XMM-Newton and Swift (the latter is not included in this table, notice), with redshift between ~ 0.2 and ~ 4.5, and X-ray luminosities, in the 0.5 - 4.5 keV band, between ~ 10<sup>43</sup> erg/s and ~ 10<sup>46</sup> erg/s. They used the structure function (SF), which operates in the time domain, and allows for an ensemble analysis even when only a few observations are available for individual sources and the power spectral density (PSD) cannot be derived. The SF is also more appropriate than fractional variability and excess variance, because these parameters are biased by the duration of the monitoring time interval in the rest-frame, and therefore by cosmological time dilation. The authors find statistically consistent results for the two samples, with the SF described by a power law of the time lag tau, approximately as SF ~ tau<sup>0.1</sup>. They do not find evidence of the break in the SF, at variance with the case of lower luminosity AGNs. They confirm a strong anti-correlation of the variability with X-ray luminosity, accompanied by a change of the slope of the SF. They also find evidence in support of a weak, intrinsic, average increase of X-ray variability with redshift. For XMM, the authors used the version of the Serendipitous Source Catalog then available, namely 2XMMi-DR3, the latest incremental update of the second version of the catalogue, with observations made between 2000 February 3 and 2008 October 08; all datasets were publicly available by 2009 October 31, but not all public observations are included in this catalog. The total area of the catalog fields is ~ 814 deg<sup>2</sup>, but taking account of the substantial overlaps between observations, the net sky area covered independently is ~ 504 deg<sup>2</sup>. The 2XMMi-DR3 catalogue contains 353,191 detections (above the processing likelihood threshold of 6), related to 262,902 unique X-ray sources, therefore a significant number of sources (41,979) have more than one record within the catalog. The selected sources were cross-correlated with the DR7 edition of the SDSS Quasar Catalog (Schneider et al. 2010, AJ, 139, 2360) to obtain redshifts and spectral classifications for the sources. The authors used a maximum distance of 1.5 arcseconds, corresponding to the uncertainty in the X-ray positions, resulting in 412 quasars that were observed by XMM-Newton from 2 to 25 epochs each for a total of 1376 observations. The authors refer to these sources as the XMM-Newton sample. This table was created by the HEASARC in April 2012 based on <a href="https://cdsarc.cds.unistra.fr/ftp/cats/J/A+A/536/A84">CDS Catalog J/A+A/536/A84</a> file table1.dat. This is a service provided by NASA HEASARC .
MAVEN (Multiple Instrument Ensemble) In-situ Observations, Data Product Bundle, Key Parameter (KP), 4 s Data
The MAVEN In-situ Calibrated Level 2, L2, Science Data Bundle contains selected fully calibrated Data from the Neutral Gas and Ion Mass Spectrometer, NGIMS, Instrument and the Particles and Fields Package together with Ephemeris Information. The Particle and Fields In-situ Instrument Data are from the Extreme Ultraviolet, EUV, Langmuir Probe and Waves, LPW, Antenna, Magnetometer, MAG, Solar Energetic Particles, SEP, SupraThermal And Thermal Ion Composition, STATIC, Solar Wind Electron Analyzer, SWEA, and Solar Wind Ion Analyzer, SWIA, Instruments. All of these Data are in Physical Units and are averaged and/or sampled at a uniform 4 s Cadence. These In-situ Data are derived directly from the Level 2 Data. The Ephemeris Data are derived by using SPICE Libraries and Kernels provided by MAVEN Navigation, NAV, Team and Lockheed-Martin.
Direct RNA sequencing and signal alignment reveal RNA structure ensembles in a eukaryotic cell
GEO Series GSE303560. Candida albicans SC5314. 20 samples. Type: Expression profiling by high throughput sequencing.
Learning-Associated Astrocyte Ensembles Regulate Memory Recall
GEO Series GSE254016. Mus musculus. 6 samples. Type: Expression profiling by high throughput sequencing.
Corticostriatal cocaine-seeking ensembles are defined by differing gene expression from sucrose-seeking ensembles using a within-subject dual self-administration and seeking mouse model
GEO Series GSE247029. Mus musculus. 60 samples. Type: Expression profiling by high throughput sequencing.
Transcript-indexed ATAC-seq reveals paired single-cell T cell receptor identity and chromatin accessibility for precision immune profiling [Ensembl]
GEO Series GSE107223. Homo sapiens. 16 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Grouped ensemble coding of behavioral states by combinations of molecularly defined cell types
GEO Series GSE148568. Mus musculus. 10 samples. Type: Expression profiling by high throughput sequencing.
SHAPE datasets from: Cooperativity in RNA chemical probing experiments modulates 2D structural ensembles
<p>Normalized SHAPE reactivity for pre-miR20b and CDE2GG, HIV-1 RRE, and FIRRE as reported in 'Cooperativity in RNA chemical probing experiments modulates RNA 2D structural ensembles'.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.