Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,577
datasets available to search
ShareScore release 0.9.0
Dataset results
2,577 results for “Inference”
Accelerating Performance Inference over Closed Systems by Asymptotic Methods
<p>This archive includes the research data associated to the paper:</p> <p>Giuliano Casale. Accelerating Performance Inference over Closed Systems by Asymptotic Methods. Proc. ACM Meas. Anal. Comput. Syst., 1(1), 2017. The paper is accepted for presentation at ACM SIGMETRICS 2017.</p> <p>The research data requires MATLAB 2015a or later. Four datasets are included, each corresponding to a section of the paper:<br> - sec5.3.1: Small and medium models without infinite server nodes (Section 5.3.1)<br> - sec5.3.2: Large models without infinite server nodes (Section 5.3.2)<br> - sec5.3.3: Models with infinite server nodes (Section 5.3.3)<br> - sec5.4: Optimization programs (Section 5.4)</p> <p>A description of each dataset is included in the README.TXT file inside each folder.</p>
A global gridded CO2 flux dataset inferred from OCO-2 retrievals using the GONGGA inversion system (v2025)
<p><strong>Data Description</strong></p> <p>Here we provide a global monthly CO2 flux dataset at 1° × 1° spatial resolution for the period 2014.9-2024.12. The dataset is generated using the GONGGA (Global ObservatioN-based system for monitoring Greenhouse GAs) inversion system by assimilating OCO-2 (Observing Carbon Observatory 2) v11.2r column CO2 retrievals that scaled to the WMO X2019 standard. The dataset contains fluxes from biosphere (Net Ecosystem Exchange, NEE) (both prior and posterior), ocean (both prior and posterior), biomass burning emissions and fossil fuel emissions.</p> <p>We also provide the posterior model simulated values corresponding to all measurements contained in the lastest release of NOAA’s ObsPack database (obspack_co2_1_GLOBALVIEWplus_v10.1_2024-11-13 and obspack_co2_1_NRT_v10.1_2025-02-07).</p> <p><strong>Change from v2024</strong></p> <ul> <li>Assimilation of OCO-2 v11.2r retrievals</li> <li>Update of prior fluxes</li> </ul> <p><strong>Data version specification</strong></p> <p>v202x.ori refers to original GONGGA flux data with 3-hourly time resolution and 2° latitude × 2.5° longitude spatial resolution, v202x refers to GONGGA flux data resampled to monthly time resolution and 1° latitude × 1° longitude spatial resolution for facilitating comparisons with other GCP inversion results.</p> <p><strong>Article citation</strong></p> <p>Jin, Z., Wang, T., Zhang, H., Wang, Y., Ding, J., Tian, X., Constraint of satellite CO2 retrieval on the global carbon cycle from a Chinese atmospheric inversion system. Science China Earth Sciences, 2023, 66: 609-618, doi: 10.1007/s11430-022-1036-7.</p> <p>Jin, Z., Tian, X., Wang, Y., Zhang, H., Zhao, M., Wang, T., Ding, J., and Piao, S.: A global surface CO2 flux dataset (2015–2022) inferred from OCO-2 retrievals using the GONGGA inversion system, Earth System Science Data, 2024, 16: 2857-2876, doi: 10.5194/essd-16-2857-2024.</p>
Data archive: CICT for single cell RNA-seq network inference
<p>This archive contains benchmarking input data and results for using single cell gene expression data to infer gene regulatory networks (GRN) by the Causal Inference with Composition of Transactions (CICT) method and a selected set of published methods. This accompanies the manuscript "Robust discovery of gene regulatory networks from single-cell gene expression data by Causal Inference Using Composition of Transactions" (Shojaee and Huang, Brief in Bioinform 2023. DOI: 10.1093/bib/bbad370). The CICT code is available at the GitHub repo (https://github.com/hlab1/scRNAseqWithCICT/).</p><p>The original CICT algorithm was described in Shojaee et al. (arXiv:1608.02658, 2016). The benchmarked methods were included in the BEELINE benchmarking pipeline (Pratapa et al., Nat Methods 2020), to which we added DEEPDRIM (Chen et al., Brief Bioinform 2021), SCENIC (Aibar et al., Nat Methods 2017), Inferelator 3.0 (Gibbs et al., Bioinformatics 2022), and CellOracle (Kamimoto et al., Nature 2023). The output directory names are (subdirectories within each dataset):</p><p>* CICT_ewMIshrink_RFmaxdepth10_RFntrees20/: CICT for simulated data<br>* CICT_v2/: CICT for experimental data<br>* CELLORACLEDB/: CellOracle for experimental data<br>* DEEPDRIM72_ewMIshrink_RFmaxdepth10_RFntrees20/: DEEPDRIM for simulated data<br>* DEEPDRIM72_v2/: DEEPDRIM for experimental data<br>* INFERELATOR38_ewMIshrink_RFmaxdepth10_RFntrees20/: Inferelator-Prior for simulated data<br>* INFERELATOR38_v2/: Inferelator-Prior for experimental data<br>* INFERELATOR34_ewMIshrink_RFmaxdepth10_RFntrees20/: Inferelator-NoPrior for experimental data<br>* INFERELATOR34_v2/: Inferelator-NoPrior for experimental data<br>* GENIE3/: GENIE3<br>* GRNBOOST2/: GRNBOST2<br>* LEAP/: LEAP<br>* PIDC/: PIDC<br>* PPCOR/: PPCOR<br>* SCENICDB/: SCENIC for experimental data<br>* SCNS/: SCNS<br>* SCODE/: SCODE<br>* SCRIBE/: SCRIBE<br>* SINCERITIES/: SINCERITIES<br>* SINGE/: SINGE<br>* RANDOM/: RANDOM</p><p>The methods were benchmarked against two kinds of scRNA-seq datasets:<br>* Simulated datasets produced by the SERGIO simulator from a synthetic network (Dibaeinia et al., Cell Systems 2020), including complete datasets and datasets with dropouts with shape parameter k=6.5 and rate parameter q=10, 30, 50, 70, 80. <br>* Experimental datasets compiled by the BEELINE pipeline, evaluated at three different levels L0, L1 and L2, with three types of ground truth networks.<br> * Evaluation levels:<br> * L0: 500 highly varying genes plus TFs<br> * L1: 1000 highly varying genes plus TFs<br> * L2: 500 highly varying genes, TFs and 500 genes randomly selected that excluded the 1000 highly varying genes from L1.<br> * Types of ground truths:<br> * Cell-type-specific ChIP-seq ground truth (L0, L1, L2)<br> * Non-specific ChIP-seq ground truth (L0_ns, L1_ns, L2_ns)<br> * Loss-of-function/gain-of-function ground truth (L0_lofgof, L1_lofgof, L2_lofgof)</p><p>The directory structure is organized in accordance with the BEELINE benchmarking pipeline. For complete details please please see the BEELINE documentation (https://murali-group.github.io/Beeline/) and Github repo (https://github.com/Murali-group/Beeline).</p><p> </p>
Raw data of healthy children and adults in a visual learning task, a conditional learning task and a transitive inference task.
<p>Raw data of 71<strong> </strong>healthy children (31 girls; average age: 6.42 years; range: 2.95-11.64 years at the beginning of the study) in a visual learning task, a 3-item conditional learning task, and a 5-item conditional learning and transitive inference task.</p><p>Raw data of 22 healthy adults (11 femaleswomen; average age: 26.05 years; range: 20.32-29.76 years at the beginning of the study) in a visual learning task, a 3-item conditional learning task, and a 5-item conditional learning and transitive inference task.</p>
Replication package and appendixes for Causal inference of server- and client-side code smells in web apps evolution
<p>-Analysis <br>--R scripts used to make the analisys, divided by folders<br>--Data folders used in the questions</p> <p>-Appendixes - used in the article to shwo extra tables and plots</p> <p>-data folders - Aggregation of data, each app has two files, CSV and xls</p> <p>-separated data folders - 5 files for each app, with lines corresponding to the each released official version<br>--serversmells<br>--clientsmells<br>--javascriptsmells<br>--Cloc(metrics)<br>--version (all oficial releases)</p> <p>-issues_bugs<br>--data -issues by app by release <br>--data_bugs_more - the same but only bugs, by app by release<br>--scripts - scrips used to aggregate issues (from daily issues to by release) anf the same for bugs</p> <p> </p>
Example inference dataset for the Ai2 Climate Emulator (ACE)
<h1>Dataset for Ai2 Climate Emulator</h1> <p> </p> <div>This dataset contains a minimal example set of files and configuration to use for inference with the Ai2 Climate Emulator (ACE). Please see https://github.com/ai2cm/ace to install the necessary software. The included checkpoint is the same ace checkpoint as referenced in (https://zenodo.org/records/10791087). See README.md for a description of the included files.<br><br>v1.1: The initial condition zarr stores were erroneously missing all values in the first upload. These files have been fixed in this update. </div> <div> </div>
Supplement to "Proof of concept for Bayesian inference of dynamic rating curve uncertainty" (v3)
<div>This deposit contains part of the updated supplement to “Proof of concept for Bayesian inference of dynamic rating curve uncertainty” (<a href="https://www.tandfonline.com/doi/full/10.1080/02626667.2024.2401094" target="_blank" rel="noopener">Cornelio et al. 2024, HSJ</a>). This version, in particular, contains two files in which the following changes were made from the earlier version (v2.0.1):</div> <div> <ul> <li><strong><em>250117_Lbn_RC_new.R</em></strong> is the updated R code. The argument for the random number generator (RNG) kind is defined for the set.seed() functions used in the script. </li> <li><strong><em>Lbn-DMs-csv0.csv</em></strong> is the updated input file containing the stage-discharge gaugings. The column for the stage values has been renamed to "H_rec" (instead of "H_m" as in the original CSV) to be consistent with the attribute name used throughout the R code.</li> </ul> <p>Except for the above files, all the input and output files in <a href="https://zenodo.org/records/12792513" target="_blank" rel="noopener">v2.0.1</a><span> remain unchanged. </span></p> </div> <p><u> </u></p>
Wikipedia: Wikipedia English - traits (inferred records)
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p>
Data for article: A quantitative framework to infer the effect of traits, diversity and environment on dispersal and extinction rates from fossils
<p>Supplementary information for:</p> <p><strong>A quantitative framework to infer the effect of traits, diversity and environment on dispersal and extinction rates from fossils</strong></p> <p>Torsten Hauffe, Mathias M. Pires, Tiago B. Quental, Thomas Wilke, and Daniele Silvestro</p> <p> </p><ul> <li> Simulations <ul> <li>Scripts <ul> <li>Scenario1_SamplingHeterogeneity.R: Script to simulate biogeographic histories with sampling heterogeneity</li> <li>Scenario3_SealevelInvasion.R: Script to simulate biogeographic histories where sea level facilitates dispersal and invasion induces extinction</li> <li>Scenario3_DiversityDependence.R: Script to simulate diversity-dependent biogeographic histories</li> <li>Scenario4_TraitDependence.R: Script to simulate trait-dependent biogeographic histories</li> <li>Scenario5_CategoricalTraitDependence.R: Script to simulate trait-dependent biogeographic histories</li> </ul> </li> <li>Results <ul> <li>Scenario1_SamplingHeterogenetiy_alpha05.txt: Results of simulations scenario 1 with a sampling heterogeneity of alpha = 0.5</li> <li>Scenario1_SamplingHeterogenetiy_alpha1.txt: Results of simulations scenario 1 with a sampling heterogeneity of alpha = 1</li> <li>Scenario1_SamplingHeterogenetiy_alpha2.txt: Results of simulations scenario 1 with a sampling heterogeneity of alpha = 2</li> <li>Scenario1_SamplingHeterogenetiy_alpha10.txt: Results of simulations scenario 1 with a sampling heterogeneity of alpha = 10</li> <li>Scenario2_Independent_dispersal_and_extinction.txt: Results of simulation scenario 2 with sea-level independent dispersal and no invasion induced extinction</li> <li>Scenario2_Sealevel_dependent_dispersal_and_independent_extinction.txt: Results of simulation scenario 2 with sea-level dependent dispersal and no invasion induced extinction</li> <li>Scenario2_Sealevel_independent_dispersal_and_invasion_induced_extinction.txt: Results of simulation scenario 2 with sea-level independent dispersal and invasion induced extinction</li> <li>Scenario2_Sealevel_dependent_dispersal_and_invasion_induced_extinction.txt: Results of simulation scenario 2 with sea-level dependent dispersal and invasion induced extinction</li> <li>Scenario3_Independent_dispersal_and_extinction.txt: Results of simulations scenario 3 with diversity-independent dispersal and extinction</li> <li>Scenario3_Diversity_dependent_dispersal_and_independent_extinction.txt: Results of simulations scenario 3 with diversity-dependent dispersal and diversity-independent extinction</li> <li>Scenario3_Independent_dispersal_and_Diversity_dependent_extinction.txt: Results of simulations scenario 3 with diversity-dependent dispersal and diversity-independent extinction</li> <li>Scenario3_Diversity_dependent_dispersal_and_extinction.txt: Results of simulations scenario 3 with diversity-dependent dispersal and extinction</li> <li>Scenario4_Independent_dispersal_and_extinction.txt: Results of scenario 4 with trait-independent dispersal and extinction</li> <li>Scenario4_Trait_dependent_dispersal_and_independent_extinction.txt: Results of scenario 4 with trait-dependent dispersal and independent extinction</li> <li>Scenario4_Independent_dispersal_and_trait_dependent_extinction.txt: Results of scenario 4 with independent dispersal and trait-dependent extinction</li> <li>Scenario4_trait_dependent_dispersal_and_extinction.txt: Results of scenario 4 with trait-dependent dispersal and extinction</li> <li>Scenario5_CatTrait_dependent_dispersal_and_independent_extinction.txt: Results of model 2 with categorical traits (e.g family) influence dispersal but no influence of a category-specific continuous traits</li> </ul> </li> </ul> </li> <li>Carnivora <ul> <li>BinnedOccurrence: Folder with 100 replicates of binned occurrences of max. 330 carnivoran genera throughout the Neogene</li> <li>BodyMass: Folder with 100 replicates of body mass for 330 carnivoran genera</li> <li>Sealevel: Folder with sea level through the Neogene</li> <li>Temperature: Folder with the temperature record of the Neogene</li> <li>Families: Folder with families as taxonomic proxy for phylogeny. FamilyGeneraNumeric.txt is the numeric coding used for the Bayesian analyses of carnivoran biogeography</li> </ul> </li> </ul> <p></p>
Inference products for "Finite inflation in curved space"
<p>These are the MCMC and nested sampling inference products and input files that were used to compute results for the paper <strong>"Finite inflation in cuved space"</strong> by <strong>L. T. Hergt</strong>, <strong>F. J. Agocs</strong>, <strong>W. J. Handley</strong>, <strong>M. P. Hobson</strong>, and <strong>A. N. Lasenby</strong> from 2022.</p> <p>Example plotting scripts (as <span>\(\texttt{.ipynb}\)</span> or as <span>\(\texttt{.html}\)</span> files) and figures from the paper are included to demonstrate usage.</p> <p> </p> <p>We used the following python packages for the genertion of MCMC and nested sampling chains:</p> <table> <tbody><tr> <th>Package</th> <th>Version</th> </tr> </tbody><tbody> <tr> <td>anesthetic</td> <td>2.0.0b12</td> </tr> <tr> <td>classy</td> <td>2.9.4</td> </tr> <tr> <td>cobaya</td> <td>3.0.4</td> </tr> <tr> <td>GetDist</td> <td>1.3.3</td> </tr> <tr> <td>primpy</td> <td>2.3.6</td> </tr> <tr> <td>pyoscode</td> <td>1.0.4</td> </tr> <tr> <td>pypolychord</td> <td>1.20.0</td> </tr> </tbody> </table> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p>Filename conventions:</p> <ul> <li><span>\(\texttt{mcmc}\)</span>: MCMC run</li> <li><span>\(\texttt{pcs#d####}\)</span>: PolyChord run (in synchronous mode) with <span>\(\texttt{#d}\)</span> repeats per parameter block (where <span>\(\texttt{d}\)</span> is the number of parameters in that block) and with <span>\(\texttt{####}\)</span> live points.</li> <li><span>\(\texttt{_cl_hf}\)</span>: Using Boltzmann theory code CLASS with nonlinearities code halofit.</li> <li><span>\(\texttt{_p18}\)</span>: Using Planck 2018 CMB data.</li> <li><span>\(\texttt{_TTTEEE}\)</span>: Using the high-l TTTEEE likelihood.</li> <li><span>\(\texttt{_TTTEEElite}\)</span>: Using the lite version of the high-l TTTEEE likelihood.</li> <li><span>\(\texttt{_lowl_lowE}\)</span>: Using the low-l likelihoods for temperature and E-modes.</li> <li><span>\(\texttt{_BK15}\)</span>: Using data from the 2015 observing season of Bicep2 and the Keck Array.</li> <li><span>\(\texttt{lcdm}\)</span>: Concordance cosmological model called LCDM (standard 6 cosmological sampling parameters, no tensor perturbations, zero spatial curvature)</li> <li><span>\(\texttt{_r}\)</span>: Extension with a variable tensor-to-scalar ratio <span>\(r\)</span>.</li> <li><span>\(\texttt{_omegak}\)</span>: Extension with a variable curvature density parameter <span>\(\Omega_K \)</span>.</li> <li><span>\(\texttt{_H0}\)</span>: Sampling over <span>\(H_0\)</span> instead of <span>\(\theta_\mathrm{s}\)</span>.</li> <li><span>\(\texttt{_omegakh2}\)</span>: Extension with a variable curvature density parameter, but sampling over <span>\(H_0\)</span> instead of <span>\(\theta_\mathrm{s}\)</span> and over <span>\(\omega_K\equiv\Omega_Kh^2\)</span> instead of <span>\(\Omega_K \)</span>.</li> <li><span>\(\texttt{_mn2}\)</span>: Using a quadratic monomial potential for the computation of the primordial universe.</li> <li><span>\(\texttt{_nat}\)</span>: Using the natural inflation potential for the computation of the primordial universe.</li> <li><span>\(\texttt{_stb}\)</span>: Using the Starobinsky potential for the computation of the primordial universe.</li> <li><span>\(\texttt{_AsfoH}\)</span>: Using the primordial sampling parameters {`logA_SR`, `N_star`, `f_i`, `omega_K`, `H0`}.</li> <li><span>\(\texttt{_perm}\)</span>: Assuming a permissive reheating scenario.</li> </ul> <p> </p> <p> </p>
MERRIN: MEtabolic Regulation Rule INference from time series data (Docker image and notebooks)
<p>This record contains notebooks and Docker image for reproducing the results of the paper "MERRIN: MEtabolic Regulation Rule INference from time series data" published as part of the ECCB 2022 conference.</p> <p>Notebooks can be executed interactively within the Docker image <code>bioasp/merrin:v1</code> which extends the <a href="http://colomoto.org/notebook">CoLoMoTo Docker</a> version <code>2021-02-01.</code></p> <p>Also see <a href="https://github.com/bioasp/merrin-covert">https://github.com/bioasp/merrin-covert</a></p> <p>The Docker image can be executed as follows:</p> <pre><code class="language-bash">docker pull bioasp/merrin:v1 docker run -it --rm -p 8888:8888 bioasp/merrin:v1 </code></pre> <p>then point your browser to <a href="http://127.0.0.1:8888">http://127.0.0.1:8888</a>.</p> <p>The image can be imported using the command <code>docker load</code> with the image file provided in this record:</p> <pre><code>docker load -i image.tar.gz</code></pre> <p>or with the <code>donodo</code> command available at <a href="https://github.com/pauleve/donodo">https://github.com/pauleve/donodo</a>:</p> <pre><code>pip install -U donodo donodo pull 10.5281/zenodo.6670165</code></pre>
Data from: Normalizing gas-chromatography–mass spectrometry data: method choice can alter biological inference
<p>Gas-Chromatography Mass Spectrometry data from European badger (<em>Meles meles</em>) sub-caudal gland secretion used in:</p> <p>Noonan, M.J., Tinnesand, H.V.,<sup> </sup>and Buesching, C.D. (2018). Normalizing gas-chromatography–mass spectrometry data: method choice can alter biological inference. BioEssays, 40(6): 0-0. DOI: 10.1002/bies.201700210.</p>
Dataset for "Sources and sinks of carbonyl sulfide inferred from atmospheric observations at the Lutjewad tower"
<p>Measurements that are used in "Sources and sinks of carbonyl sulfide inferred from atmospheric observations at the Lutjewad tower”. The dataset includes mole fraction measurements of COS, CO<sub>2</sub>, CO and H<sub>2</sub>O made in Lutjewad (the Netherlands) and Hyytälä (Finland) between 2014 and 2018, and measurements of <sup>222</sup>Rn, SF<sub>6</sub> and meteorological parameters in Lutjewad.</p>
1D cell trajectories as studied in "Cell-mechanical parameter estimation from 1D cell trajectories using simulation-based inference"
<p>Trajectories of motile cells represent a rich source of data that provide insights into the mechanisms of cell migration via mathematical modeling and statistical analysis. Here, we present trajectories of MDA-MB-231 breast cancer cells and MCF-10A breast epithelial cells. Cells were confined to 1D using fibronectin lanes and exposed to three different treatments, namely the actin polymerisation inhibitor Latrunculin A (LatA), the ROCK inhibitor Y-27632 (Y27) and a control. Each csv file contains a number of 24h long trajectories of cells corresponding to the name of the file. The column names are:</p> <p>`traject_id`: The trajectories are numbered, starting from 0 in each file.</p> <p>`time (h)`: Time in h, starting at 0h for each trajctory and ending at 24h with a temporal resolution of 2min.</p> <p>`x_front`: Position of the cell's front.</p> <p>`x_nucleus`: Position of the cell's nucleus, where x_nucleus=0 for the first time point of the trajectory</p> <p>`x_rear`: Position of the cell's rear.</p> <p>The data was analysed in our study "Cell-mechanical parameter estimation from 1D cell trajectories using simulation-based inference". Further information can be found there. </p>
CLDF dataset accompanying List's "Inference of Partial Colexifications" from 2023 building on Key and Comrie's "Intercontinental Dictionary Series" from 2021
<p>Cite the source of the dataset as:</p> <blockquote> <p>List, J.-M. (2023): Inference of partial colexifications from multilingual wordlists. Frontiers in Psychology 14.1156540. 1-10.</p> </blockquote>
Inference Ontology
<p>Overview</p> <p>The diagrams showcase an inference ontology designed to model the provenance of any proposition or statement. Although originally developed within the context of the <a href="https://github.com/GOLEM-lab/golem-ontology">GOLEM framework</a>, which has since adopted a different approach, this module remains applicable across various digital humanities domains by focusing on the inference-making process for information, utilizing the <a href="https://www.cidoc-crm.org/crminf/fm_releases">CRMinf ontology.</a></p> <p>Scope<br>The Inference Module addresses three key aspects of inference:</p> <ul> <li>Source for Inference: The origin or basis of the inference.</li> <li>Method of Inference: The logic or methodology applied to derive the inference.</li> <li>Premise for Inference: Other statements or propositions that the inference is based on.</li> </ul> <p>Ontology (Diagram)<br>In this ontology, any statement is represented as crminf:I2_Belief. The formal representation of this belief is encapsulated in crminf:I4_Proposition_Set, which can be:</p> <ul> <li>A single RDF triple (atomic proposition).</li> <li>Multiple triples, or an entire knowledge graph (composite proposition).</li> </ul> <p>The core of these statements is structured based on the subject–predicate–object framework. RDF reification is employed to treat the entire statement as the subject in the inference triple.</p> <p>Key Components</p> <ul> <li>Belief: Concluded from the inference process, represented by crminf:I5_Inference_Making.</li> <li>Source of Inference: Defined by crm:E73_Information_Object, which may be a text, metadata, or any other form of information.</li> <li>Method of Inference: Captured by crminf:I3_Inference_Logic, which can be associated with crm:E55_Type to specify methods like manual annotation or machine learning.</li> <li>Premise: Other beliefs or statements upon which the inference relies, formalized by crminf:I4_Proposition_Set.</li> </ul> <p>Classes</p> <ul> <li>crminf:I2_Belief: Represents any proposition within the ontology.</li> <li>crminf:I4_Proposition_Set: The formal representation of a belief.</li> <li>crminf:I5_Inference_Making: The process of making an inference for any proposition.</li> <li>crminf:I3_Inference_Logic: The methodology or logic used in the inference.</li> <li>crm:E73_Information_Object: The source of inference, specified by subclasses such as lrm:F2_Expression for text or</li> <li>crm:E33_Linguistic_Object for metadata.</li> <li>rdf:Statement: Each crminf:I4_Proposition_Set is also an rdf:Statement, allowing specification of the subject, predicate, and object using rdf:subject, rdf:predicate, and rdf:object.</li> </ul> <p>Properties</p> <ul> <li>crminf:J4_that: Links a belief to its formal representation.</li> <li>crminf:J2_concluded_that: Links a belief to the inference that concluded this belief.</li> <li>crminf:J3_applies: Links an inference process to its method.</li> <li>crminf:J1_used_as_premise: Links the belief used as a premise for another belief.</li> <li>crm:P16_used_specific_object: Links the source to the inference process.</li> </ul> <p>Example: Relationship Inference (Diagram)<br>This example illustrates how relationships between characters are inferred based on events.</p> <ul> <li>Statement<br>Hermione Granger and Ron Weasley are involved in a romantic relationship. <ul> <li>Subject: romantic_love (dlp:relationship)</li> <li>Predicate: dlp:involves</li> <li>Object: Hermione Granger and Ron Weasley (gc:G1_Character)</li> <li>Inference Process</li> </ul> </li> <li>The inference-making process for this relationship statement includes: <ul> <li>Premise: The relationship is based on events. In this case, the premise is the event statement: "Hermione Granger and Ron Weasley participated in a kissing event during the Battle of Hogwarts."</li> <li>Method: This relationship was inferred based on ontology rules, specifically that the relationship forms based on events.</li> <li>Premise Belief: The premise belief (the kissing event) necessitates another inference process. This secondary inference process utilizes a textual source as evidence for the kissing event, which is a paragraph in the work "Harry Potter and the Deathly Hallows".</li> </ul> </li> <li>Premise Belief Inference <ul> <li>Event Statement: "Hermione Granger and Ron Weasley participated in a kissing event during the Battle of Hogwarts."</li> <li>Source: The source for this event is a crm:E73_Information_Object, potentially a text or metadata describing the event.</li> <li>Inference Method: This event is inferred through manual annotation of the text, which is linked via crminf:I3_Inference_Logic.</li> </ul> </li> </ul>
Processed Data for "Improving Gene Regulatory Network Inference using Dropout Augmentation"
<p>Here are the processed dataset that are used in the manuscript "Improving Gene Regulatory Network Inference using Dropout Augmentation"</p>
CLDF dataset derived from Satterthwaite-Phillips' "Phylogenetic Inference of the Tibeto-Burman Languages" from 2011
<p>Cite the source of the dataset as:</p> <blockquote> <p>Satterthwaite-Phillips, Damian (2011) Phylogenetic inference of the Tibeto-Burman languages or on the usefuseful of lexicostatistics (and "megalo"-comparison) for the subgrouping of Tibeto-Burman. Stanford: Stanford University.</p> </blockquote>
A unified genealogy of modern and ancient genomes: Unified, inferred tree sequences of 1000 Genomes, Human Genome Diversity, and Simons Genome Diversity Projects
<p>Unified, inferred tree sequences built from the 1000 Genomes phase 3, Human Genome Diversity, and Simons Genome Diversity Projects. Each tree sequence is the arm of an autosome (the short arm of acrocentric chromosomes are not included). Tree sequences were inferred using <a href="https://tsinfer.readthedocs.io/">tsinfer</a> version 0.2.1, dated using <a href="https://tsdate.readthedocs.io/en/latest/">tsdate</a> version 0.1.4 and compressed using <a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. All data is in GRCh38.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on <a href="https://github.com/awohns/unified_genealogy_paper">GitHub</a>. A description can be found in the Supplementary Material of <a href="https://www.biorxiv.org/content/10.1101/2021.02.16.431497v2">Wohns et al. (2021)</a>.</p> <p>Tree sequences can be decompressed as follows:</p> <pre><code>$ tsunzip hgdp_tgp_sgdp_chr1_p.dated.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed in Python using <a href="https://tskit.readthedocs.io/">tskit</a>. </p> <pre><code>import tskit ts = tskit.load("hgdp_tgp_sgdp_chr1_p.dated.trees") # ts is an instance of tskit.TreeSequence print("The short arm of chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with nodes contain the mean and variance of tsdate's posterior distribution on node time. To access these values, we can use:</p> <pre><code>import json node = ts.node(10000) metadata_dict = json.loads(node.metadata) print("The mean of the posterior distribution on the age of node 10000 is {} generations".format(metadata_dict["mn"])) print("The variance of the posterior distribution on the age of node 10000 is {} generations".format(metadata_dict["vr"]))</code></pre> <p>Age estimates for each variant site can be derived from the mean of the age estimates of the upper and lower bounding nodes of the oldest mutation associated with a site. tsdate includes <a href="https://tsdate.readthedocs.io/en/latest/python-api.html?highlight=sites_time_from_ts#tsdate.sites_time_from_ts">a function to find the age estimates of all sites in the tree sequence</a>:</p> <pre><code>import tsdate site_times = tsdate.sites_time_from_ts(ts, node_selection='arithmetic')</code></pre> <p>This returns a numpy array which has a length equal to the number of sites.</p> <p>Accessing variant sites in the tree sequence provides the position and id of variants:</p> <pre><code>site = ts.site(1000) site_metadata = json.loads(site.metadata) print("The position of site 1000 is {} and its ID is {}.".format(site.position, site_metadata["ID"]))</code></pre> <p>Metadata associated with individuals and populations was derived from the original sources (<a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/working/20130606_sample_info/20130606_g1k.ped">TGP</a>, <a>HGDP</a>, and <a href="https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/SGDP_metadata.279public.21signedLetter.samples.txt">SGDP</a>) and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code>ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code>pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p>
Gene Regulatory Network inference in long lived C.elegans reveals modular properties that are predictive of novel ageing genes - Supplementary Tables
<p>This repository contains the Supplementary Tables for Suriyalaksh et al. Gene Regulatory Network inference in long lived C.elegans reveals modular properties that are predictive of novel ageing genes.</p> <p>The list of table files can be found in <a href="https://zenodo.org/api/files/431155b3-c2b2-4fff-92bb-e93cad8bf5c2/Supplementary%20table%20guide.pdf">Supplementary Tables guide.pdf</a></p> <p>Tables S1, S2 and S3 corresponding to physical gene-gene interaction data are in a separate repository doi:10.5281/zenodo.4382337</p> <p>Details about some of the Supplementary tables:</p> <p>TableS4_inferred_networks.csv - list of inferred GRNs for specified input combinations (set of input regulators, length of the time sequence, NI tool and prior used).</p> <p><a href="https://zenodo.org/api/files/431155b3-c2b2-4fff-92bb-e93cad8bf5c2/TableS5_consensus_network_member.xlsx">TableS5_consensus_network_member.xlsx</a> - list of groups of topologically similar GRNs (from Table S4)</p> <p>Table S6: edge lists (source,target) for each one of the three consensus networks selected according to the GS validation metrics: middle PFE/AUFE, max AUFE, max PFE.<br> TableS6a_max_AUFE_GRN.txt - max AUFE; largest network - this is the one we used in the main analysis and discussion<br> TableS6b_max_PFE_GRN.xt - max PFE<br> TableS6c_middle_AUFE_PFE_GRN.txt - middle PFE/AUFE</p> <p><a href="https://zenodo.org/api/files/431155b3-c2b2-4fff-92bb-e93cad8bf5c2/TableS7_qRTPCR_ddCt_network_accuracy.csv">TableS7_qRTPCR_ddCt_network_accuracy.csv</a> - gene expression count differences for RNAi knockdown GRN validation experiments. </p> <p>Table S8: Group membership for each one of the nodes in each one of the selected networks according to the SBM that best describes the observed network topology. Each column shows the group membership for each level in a SBM block hierarchy. Our analysis is in the second most coarse-grained level (level 1).</p> <p>TableS8a_max_AUFE_SBM.csv<br> TableS8b_max_PFE_SBM.csv<br> TableS8c_middle_AUFE_PFE_SBM.csv</p> <p><a href="https://zenodo.org/api/files/431155b3-c2b2-4fff-92bb-e93cad8bf5c2/TableS9_glp_gs_datasets.pdf">TableS9_glp_gs_datasets.pdf</a> - list of datasets used for defining functional clusters.</p> <p>TableS14a_glp_l1_vs_fem_l1_lifespan_assay.xlsx - Day13 survival of fem-3(q20)ts vs day 19 survival of glp-1(e2144)ts;rrf-3(pk1426) RNAi from L1</p> <p>TableS14b_glp_l1_vs_glp_l4_lifespan_assay.xlsx - Day 19 survival of glp-1(e2144)ts;rrf-3(pk1426) RNAi from L4 vs day 19 survival of glp-1(e2144)ts;rrf-3(pk1426) RNAi from L1</p> <p>TableS15a_glp1_in_vivo_fluorescence_data.xlsx - in vivo fluorescent reporter data of glp-1(e2144)ts;rrf-3(pk1426)</p> <p>TableS15b_fem3_in_vivo_fluorescence_data.xlsx - in vivo fluorescent reporter data of fem-3(q20)ts</p> <p>TableS17_input_regulators_annotated.csv - list input regulators used as input for Network Inference Tools annotated by source type (2nd column): GenAge, known transcription factors (TF) and gene with high variability in the gene expression time series (HV). The third column lists whether that regulator has an orthologue in human (y) according to WormBase (v 278).</p> <p>TableS20_epistasis_lifespan_data.xlsx - Epistasis lifespan data of glp-1(e2144)ts</p> <p>All the image (TIF) files represent representative images in the following genetic backgrounds (below) that have been treated </p> <p>with empty vector (EV) or RNAi against the gene highlighted in the title of the image. See methods section for details. </p> <p><strong>femliu1: </strong></p> <p><em>fem-3(q20)ts.; dhs-3p::dhs-3::gfp</em></p> <p><strong>femsod3:</strong></p> <p><em>fem-3(q20)ts.; sod-3p::gfp</em></p> <p><strong>glp1lgg1:</strong></p> <p><em>glp-1(e2144); lgg-1p:lgg-1:gfp</em></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.