Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
608
datasets available to search
ShareScore release 0.9.0
Dataset results
608 results for “ensembles”
The hydrological fluxes of the Upper Brahmaputra River Basin constrained by a multi-physics ensemble (MPE) modeling approach
<p>The data represent monthly hydrological fluxes for four sub-basins within the Upper Brahmaputra River Basin, where yyyy is the year, mm is the month, MPE is the multi-physics ensemble, P is the precipitation, R is the runoff, and ET is the evapotranspiration. The unit is mm. The upper-bounds and lower-bounds represent the upper and lower bounds of the hydrological fluxes constrained by the MPE, respectively.</p> <p> </p> <p>Reference:<br>Lei, X, P. Lin*, H. Zheng, K. Yang, W. Liu, C. Miao, K. Wang, J. Wang: A multi-physics ensemble modeling approach to constraining the uncertainty of hydrological fluxes in sparsely-gauged river basins. Geophysical Research Letters, (submitted), 2024.<br>Contact:<br>xiangyonglei@stu.pku.edu.cn; peironglinlin@pku.edu.cn</p>
nextsim ensemble perturbation test data
Open the record for dataset details and reuse information.
GeneSeqToFamily: a Galaxy workflow to find gene families based on the Ensembl Compara GeneTrees Pipeline.
<p>Gene duplication is a major factor contributing to evolutionary novelty, and the contraction or expansion of gene families has often been associated with morphological, physiological and environmental adaptations. The study of homologous genes helps us to understand the evolution of gene families. It plays a vital role in finding ancestral gene duplication events as well as identifying genes that have diverged from a common ancestor under positive selection. There are various tools available, such as MSOAR, OrthoMCL and HomoloGene, to identify gene families and visualise syntenic information between species, providing an overview of syntenic regions evolution at the family level. Unfortunately, none of them provide information about structural changes within genes, such as the conservation of ancestral exon boundaries amongst multiple genomes. The Ensembl GeneTrees computational pipeline generates gene trees based on coding sequences and provides details about exon conservation, and is used in the Ensembl Compara project to discover gene families. </p>
Ensembl genes [data set for teaching purposes]
<p>CSV file for teaching purposes.</p>
Data from: Detecting epistasis from an ensemble of adapting populations
The role that epistasis plays during adaptation remains an outstanding problem, which has received considerable attention in recent years. Most of the recent empirical studies are based on ensembles of replicate populations that adapt in a fixed, laboratory controlled condition. Researchers often seek to infer the presence and form of epistasis in the fitness landscape from the time-evolution of various statistics averaged across the ensemble of populations. Here we provide a rigorous analysis of what quantities, drawn from time-series of such ensembles, can be used to infer epistasis for populations evolving under weak mutation on finite-site fitness landscapes. First we analyze the mean fitness trajectory—that is, the time course of the ensemble average fitness. We show that for any epistatic fitness landscape and starting genotype, there always exists a non-epistatic fitness landscape that produces the exact same mean fitness trajectory. Thus, the presence of epistasis is not identifiable from the mean fitness trajectory. By contrast, we show that two other ensemble statistics—the time evolution of the fitness variance across populations, and the time evolution of the mean number of substitutions—can detect certain forms of epistasis in the underlying fitness landscape.
An ensemble deep learning framework to refine large deletions in linked-reads
<p>An ensemble deep learning framework to refine large deletions in linked-reads</p>
Atmospherically-forced and chaotic interannual variability of the sea level and its components over 1993-2015 from the OCCIPUT ensemble simulations
<p>This data set contains the interannual variability fields for the sea level (ssh_var_inter_1993_2015_annuel.tar.gz) and its steric (hsterica_var_inter_1993_2015_annuel.tar.gz) and manometric (obp_var_inter_1993_2015_annuel.tar.gz) components over 1993-2015 from the OCCIPUT ensemble simulations. It is used in the paper « Atmospherically forced and chaotic interannual variability of regional sea level and its components over 1993-2015 » published in Journal of Geophysical Research - Oceans.</p> <p>This dataset has been computed from the OceaniC Chaos – ImPacts, strUcture, predicTability (OCCIPUT) global ocean/sea-ice ensemble simulation. It is composed of 50 members with a horizontal resolution of 1/4° and 75 geopotential levels (Bessières et al., 2017, Penduff et al., 2014). The numerical configuration is based on the version 3.5 of the NEMO model (Madec, 2008). The 50 members were started on January 1st 1960 from a common 21-year spinup. A small stochastic perturbation is applied to the equation of state of sea water (as in Brankart, 2013) within each member during 1960, then switched off during the rest of the simulation. This 1-year perturbation generates an ensemble spread which grows and saturates after a few months up to a few years depending on the region. The 50 members are driven through bulk formulae during the whole 1960-2015 simulation by the same realistic 6-hourly atmospheric forcing (Drakkar Forcing Set DFS5.2, Dussin et al., 2016) derived from ERA interim atmospheric reanalysis.</p> <p> </p> <p>For each member, the simulated sea surface height (SSH) over 1993-2015 is considered. As NEMO is a Boussinesq model, it conserves volume instead of mass. Therefore, the steric effect is missing into the global mean sea level change (Greatbatch 1994). To overcome this issue, we remove the global mean estimate for the sea level time series at each grid point. Then the sea level anomalies obtained are averaged per year and a linear trend is removed from each member. The same processes are applied to the steric and manometric sea level time series.</p> <p> </p> <p>Here is an example of the file header</p> <p><em><strong>dimensions:</strong></em></p> <p><em><strong>member = UNLIMITED ; // (50 currently)</strong></em></p> <p><em><strong>time = 23 ;</strong></em></p> <p><em><strong>y = 1021 ;</strong></em></p> <p><em><strong>x = 1442 ;</strong></em></p> <p><em><strong>variables:</strong></em></p> <p><em><strong> float ssh(member, time, y, x) ;</strong></em></p> <p><em><strong> ssh:long_name = "sea level interannual variability" ;</strong></em></p> <p><em><strong> ssh:standard_name = "sea_level interannual variability" ;</strong></em></p> <p><em><strong> ssh:units = "m" ;</strong></em></p> <p><em><strong> ssh:FillValue = "nan" ;</strong></em></p> <p><em><strong> float nav_lat(member, y, x) ;</strong></em></p> <p><em><strong> nav_lat:axis = "Y" ;</strong></em></p> <p><em><strong> nav_lat:long_name = "Latitude" ;</strong></em></p> <p><em><strong> nav_lat:standard_name = "latitude" ;</strong></em></p> <p><em><strong> nav_lat:units = "degrees_north" ;</strong></em></p> <p><em><strong> float nav_lon(member, y, x) ;</strong></em></p> <p><em><strong> nav_lon:axis = "Y" ;</strong></em></p> <p><em><strong> nav_lon:long_name = "Longitude" ;</strong></em></p> <p><em><strong> nav_lon:standard_name = "longitude" ;</strong></em></p> <p><em><strong> nav_lon:units = "degrees_east" ;</strong></em></p> <p><em><strong> float time(time) ;</strong></em></p> <p><em><strong> time:long_name = "time" ;</strong></em></p> <p><em><strong> time:standard_name = "time" ;</strong></em></p> <p><em><strong> time:units = "years since 1992" ;</strong></em></p> <p> </p> <p>nav_lat and nav_lon represent the latitude and longitude of the NEMO model whereas var represents the interannual variability time series. </p>
Figure 2 from: Vyshedskiy A, Dunn R (2015) Mental synthesis involves the synchronization of independent neuronal ensembles. Research Ideas and Outcomes 1: e7642. https://doi.org/10.3897/rio.1.e7642
Figure 2 - On a neurological level, mentally forming the image of Bill Clinton and the lion consists of the following steps: Step 1 - Recall of Bill Clinton: The prefrontal cortex (PFC) activates the ensemble of neurons representing Bill Clinton to fire synchronous actions potentials. Bill Clinton is perceived by the patient. The electrode implanted into the temporal lobe (TL) records an increased rate of action potentials. Step 2 - Recall of the lion: The PFC activates the ensemble of neurons representing the lion to fire synchronous actions potentials. The lion is perceived. The second electrode implanted into the TL records an increased rate of action potentials. Step 3 - The patient mentally integrates the images of Bill Clinton and the lion into one scene. The Mental Synthesis Theory hypothesizes that integration is accomplished by the PFC synchronizing the two neuronal ensembles in time. Step 4 - When synchronization of the Clinton and the lion neuronal ensembles is achieved, a new, never-before-seen mental image of Bill Clinton holding the lion on his lap is perceived by the patient. At that moment the two implanted electrodes are predicted to record synchronous action potentials, implying the synchronization of the Clinton and lion neuronal ensembles.
Figure 1 from: Vyshedskiy A, Dunn R (2015) Mental synthesis involves the synchronization of independent neuronal ensembles. Research Ideas and Outcomes 1: e7642. https://doi.org/10.3897/rio.1.e7642
Figure 1 - Mental synthesis of Bill Clinton holding a lion. Once selective neurons for Bill Clinton and the lion are identified, a subject can be asked to imagine Bill Clinton holding the lion on his lap. The Mental Synthesis theory predicts that both the Clinton neuron and the lion neuron will increase their firing rate and that their activity will be synchronized. Images modified from: 1. William J. Clinton at the Parliament in London, United Kingdom, November 29, 1995. https://commons.wikimedia.org/wiki/File:Bill_Clinton_1995_im_Parlament_in_London.jpg 2. Lioness in the Olomouc Zoo at Svatý kopeček, Czech Republic. This image is licensed under the CC BY-SA license. https://commons.wikimedia.org/wiki/File:Lioness,_Olomouc.jpg
FIRO_synthetic-ensemble-forecasts dataset v0
<p>Dataset to support initial submission of WRR manuscript 'Synthetic forecast ensembles for evaluating Forecast Informed Reservoir Operations (FIRO)'</p> <p>Code and workflow description are in public GitHub repository: https://github.com/zpb4/FIRO_synthetic-ensemble-forecasts</p>
LENS and NoPin (NP) ensemble means for apparent oxygen utilization (AOU), anthropogenic and preindustrial carbon, pCFC-12, and ideal age from Olivarez et al. (2023)
<p>Manuscript title: <strong>"How does the Pinatubo eruption influence our understanding of long-term changes in ocean biogeochemistry?"</strong></p> <p>Authors: Holly C. Olivarez<sup>1,2</sup>, Nicole S. Lovenduski<sup>2,3</sup>, Yassir A. Eddebbar<sup>4</sup>, Amanda R. Fay<sup>5</sup>, Galen A. McKinley<sup>5</sup>, Michael N. Levy<sup>6</sup>, and Matthew C. Long<sup>6</sup></p> <p>Affiliations:<br> <sup>1</sup>Department of Environmental Studies, University of Colorado, Boulder, CO, USA <br> <sup>2</sup>Institute of Arctic and Alpine Research, University of Colorado, Boulder, CO, USA<br> <sup>3</sup>Department of Atmospheric and Oceanic Sciences, University of Colorado, Boulder, CO, USA <br> <sup>4</sup>Scripps Institution of Oceanography, University of California San Diego, La Jolla, CA, USA <br> <sup>5</sup>Columbia University and Lamont-Doherty Earth Observatory, Palisades, NY, USA <br> <sup>6</sup>Climate and Global Dynamics Laboratory, National Center for Atmospheric Research, Boulder, Colorado, USA</p> <p>LENS and No Pinatubo (NoPin in manuscript, NP in file names) CESM ensemble means for apparent oxygen utilization (AOU), anthropogenic and preindustrial carbon, <em>p</em>CFC-12, and ideal age along GO-SHIP transects akin to <a href="https://cchdo.ucsd.edu/data/4925/33RO20130803do.pdf">A16N</a>, <a href="https://cchdo.ucsd.edu/data/9276/33RO20131223do.pdf">A16S</a>, <a href="https://cchdo.ucsd.edu/data/14283/33RO20150525_do.pdf">P16N</a>, <a href="https://cchdo.ucsd.edu/cruise/320620140320#note_334284">P16S</a>, <a href="https://cchdo.ucsd.edu/cruise/09AR20060102">SO4I</a>, <a href="https://cchdo.ucsd.edu/cruise/320620180309">SO4P</a>.</p> <p>The CESM source code is freely available at <a href="http://www2.cesm.ucar.edu">http://www2.cesm.ucar.edu</a>. The model outputs described in this paper can be accessed at <a href="http://www.earthsystemgrid.org">www.earthsystemgrid.org</a>.</p>
Ensemble of initial conditions for Hydro-ABC
Open the record for dataset details and reuse information.
Novel ensemble models for groundwater potential mapping: Application of the Split-Point Sampling and Node Attribute Subsampling Classifier in Vietnam
<p>Le Tien Duy</p>
Development of an Ensemble Learning-based, Multi-dimensional Sensory Impairment Score to Predict Cognitive Impairment in an Elderly Cohort of Southern Italy
ClinicalTrials.gov study NCT06783010. IPD Sharing: UNDECIDED. Countries: 0. Publications: 8.
Data from: Comparing radiomic classifiers and classifier ensembles for detection of peripheral zone prostate tumors on T2-weighted MRI: a multi-site study
Open the record for dataset details and reuse information.
Data from: Detecting epistasis from an ensemble of adapting populations
Open the record for dataset details and reuse information.
Data from: Ensembler: enabling high-throughput molecular simulations at the superfamily scale
Open the record for dataset details and reuse information.
Assembling ensembling: An adventure in approaches across disciplines
Open the record for dataset details and reuse information.
Data from: Ensemble modelling of the potential distribution of the whale shark in the Atlantic Ocean
Open the record for dataset details and reuse information.
NASA MEASURES Precipitation Ensemble based on SSMIS DMSP F18 NASA PPS L1C V05 Tbs 1-orbit L2 Swath 12x12km V1 (PRECIP_SSMIS_F18) at GES DISC
The data presented in this level 2 orbital product are rain rate estimates expressed as mm/hour determined from brightness temperatures (Tbs) obtained from the Special Sensor Microwave Imager Sounder (SSMIS) flown on the US Defense Meteorological Satellite Program (DMSP) F18 mission. Most of the products generated in this data set are based upon the algorithms developed for the 3rd Algorithm Intercomparison Project (AIP-3) of the Global Precipitation Climatology Project (GPCP). Details of these 15 algorithms and development of a quality score which is a measure of confidence in the estimate, along with processing and algorithmic flags, can be found in the Algorithm Theoretical Basis Document (ATBD). The data in this product cover the period from 2010 to 2020 with one file per orbit.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.