Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
492
datasets available to search
ShareScore release 0.9.0
Dataset results
492 results for “sequence modeling”
Emergenet models for HA/NA sequences of Influenza A strains
<p>Models emergent dependencies between mutational variations of key viral proteins in Influenza A.</p> <p>Models can be read using the emergenet package available on pypi. (https://pypi.org/project/emergenet/)</p>
Genome report: Genome sequence of 1S1, a transformable and highly regenerable diploid potato for use as a model for gene editing and genetic engineering
<p>Generation of a genomic resource for a readily transformable diploid potato would provide a resource for high throughput functional analysis in potato. The heterozygous <em>Solanum tuberosum</em> Group Phureja clone 1S1 has a high regeneration rate, self-fertility, desirable tuber traits and is amenable to <em>Agrobacterium</em>-mediated transformation. To create a contiguous genome assembly, a homozygous doubled monoploid of 1S1 (DM1S1) was sequenced using 44 Gbp of long reads generated from Oxford Nanopore Technologies (ONT), yielding a 736 Mb assembly that encoded 31,145 protein-coding genes. The final assembly for DM1S1 represents a nearly complete genic space, shown by the presence of 99.6% (C:99.5%[S:97.8%, D:1.7%],F:0.1%,M:0.4%,n:1614) of the Benchmarking Universal Single Copy Orthologs. Variant analysis with Illumina reads from 1S1 was used to deduce its alternate haplotype using the variant calling tools Strelka2 (v2.9.10), GATK's Haplotypecaller (v4.1.4.1), and Freebayes (v1.3.2). These variants were used to create consensus fasta sequences with the DM1S1 assembly using bcftools (v1.9.64).</p>
Datasets for manuscript - Dirichlet diffusion score model for biological sequence generation.
<p>This repository holds the trained Dirichlet Diffusion Score models for various datasets.</p> <p><strong>best_models.tar.gz</strong></p> <p>It also contains all input data required to train your own models with scripts provided via <a href="https://github.com/jzhoulab/ddsm">github repository</a>.</p> <p><strong>data.tar.gz</strong></p> <p>This archive contains the following folders: </p> <ul> <li><strong>satnet_sudoku </strong>contains dataset with sudoku examples which we used for evaluation of sudoku model.</li> <li><strong>promoter_design</strong> contains dataset used for training promoter design model as well as Sei model weights. Please, read provided readme file before using it for training scripts. </li> </ul>
Model-based analysis of sample index hopping reveals its widespread artifacts in multiplexed single-cell RNA-sequencing
<p>Supplementary data that are needed to rerun the reproducible notebooks from the first steps using Alevin output and configuration files.</p> <p>Intermediate R data object that can be used to rerun the reproducible notebooks after the filtering steps.</p> <p>Validation data for inferring the sample index hopping rate. The <em>hiseq4000_joined_datatable_plexed_nonplexed.zip file contains read counts for four samples (two non-multiplexed and two multiplexed) joined by a cell-barcode, UMI, and gene-ID (CUG) key combination. The hiseq4000_inner_joined_with_labels.zip file contains only those CUGs that are observed in both the non-multiplexed and multiplexed samples.</em><em> </em></p>
Geodetic dataset and slip models of the 2022 Chihshang earthquake sequence in eastern Taiwan
<p>This repository contains the geodetic dataset and corresponding slip models presented in Tang et al. (2023), "Nearby fault interaction within the double-vergence suture in eastern Taiwan during the 2022 Chihshang earthquake sequence".</p>
Automatic message sequence chart creation from simulation run of the Chandy-Lamport algorithm modeled by colored Petri net
<p>Videos of two message sequence chart creation from simulation runs of the Chandy-Lamport algorithm modeled by colored Petri net using the CPN tool.</p> <p><strong>Message Sequence Chart Of Model With Automatic Simulation Run_SuppInfo.mp4</strong>: This video shows the automatic generation of a message sequence chart of a proposed colored Petri net model of the Chandy-Lamport distributed global snapshot algorithm using the CPN tool version 4.0.1. The video has been generated by the authors' updated extension server of the CPN tool. The automatic simulation run of the model has been used to create this video. The CPN tool randomly selects the enabled transition in an automatic simulation run.</p> <p><strong>Message Sequence Chart Of Model With Step-By-Step Simulation Run_SuppInfo.mp4:</strong> This video shows the automatic generation of a message sequence chart of the proposed colored Petri net model of the Chandy-Lamport algorithm in a step-by-step simulation run with our updated extension server of the CPN tools version 4.0.1. We fired our selected enabled transition of the model to create this video.</p>
Data from: A cost-effective blood DNA methylation-based age estimation method in domestic cats, Tsushima leopard cats (Prionailurus bengalensis euptilurus), and Panthera species, using targeted bisulfite sequencing and machine learning models
Open the record for dataset details and reuse information.
Genome report: Genome sequence of 1S1, a transformable and highly regenerable diploid potato for use as a model for gene editing and genetic engineering
Open the record for dataset details and reuse information.
Complex models of sequence evolution improve fit, but not gene tree discordance, for tetrapod mitogenomes
Open the record for dataset details and reuse information.
Branching Trees from standard Epidemic Aftershock Sequences (ETAS) Model (no time, no space)
<p>All files licensed under Creative Commons Attribution 4.0 International (CC BY 4.0)</p> <p>###############<br> 0. SUMMARY<br> ###############</p> <p>1. DESCRIPTION</p> <p>2. INPUT PARAMETERS</p> <p>3. TYPES OF FILES<br> 3.1. Raw data<br> 3.2. List of trees<br> 3.3. Tree-size frequencies</p> <p>4. LIST OF FILES<br> 4.1. Raw data<br> 4.2. List of trees<br> 4.3. Tree-size frequencies<br> 4.4 Known missing/broken files</p> <p>###############<br> 1. DESCRIPTION<br> ################<br> Simulation results of an standard ETAS model as a branching process. Using a two seed version of the RANDU linear congruential pseudorandom number generator. The offspring number is a Poisson number given the rate n(M) (see below). Details of simulation procedure can be found in reference [1]: 'Topological properties of epidemic aftershock processes', by J. Baró submitted to J. of Geophysical Research - Solid Earth (JGR-B)</p> <p>##############<br> 2. INPUT PARAMETERS<br> ##############<br> The input parameters (see reference for details) for each raw and processes data-file are indicated in the prefix of the file: "ETASbranch_b(b)r(r)N(nb)*"</p> <p>- M0(= 1) = magnitude of completeness (arbitrary for the study of topological properties of trees)<br> - (b) = b-value (arbitrary for the study of topological properties of trees)<br> - (nb) = average branching ratio<br> - (r) = ratio a/b</p> <p>The b-value defines the distribution of event-magnitudes: P(M) = 10^(b*(M-M0)) . The nb and a define the productivity law: n(M) = (nb*(b-a)/b)*10^(a*(M-M0))</p> <p><br> #################<br> 3. TYPES OF FILES<br> #################</p> <p>3.1. Raw data:<br> --------------</p> <p>42 x "*.Seq" files with input b=0.50 and different nb, r values. Raw data from simulation code. (all cases, simulated with 10^5 background events)<br> Each row represents an individual event in the point process, or element of the simulated branching forest.<br> Columns description (9 columns x data point):<br> c0:Time (arbitrary, used here as id.)<br> c1:Magnitude of the event<br> c2:Identification number of the cluster or tree<br> c3:Depth of the event in the tree structure (background events have Depth = 0)<br> c4:Time of the direct parent of the event (set to -1 for background event)<br> c5:Magnitude of the direct parent of the event (set to -10 if background event)<br> c6:Time of the event initiating the tree (set to own time if background)<br> c7:Magnitude of the event initiating the tree (set to own magnitude if background)<br> c8:N or offspring number of the event. (Events are leafs if N=0)</p> <p><br> 3.2. List of trees<br> ------------------</p> <p>53 x "*TopoTrees.dat" files obtained from simulations (after processing of *.Seq files. Files ending with "N0.99", "N0.50" obtained from 10^5 background events from files above. Files ending with "N0.500*" obtained from 10^7 simulations)<br> Each row represents an individual tree constituted by one or several causally connected events of the simulated branching forest. Files used to generate fig. 4 of ref. [1]</p> <p>Columns description (8 columns x data point):<br> c0:Maximum Depth of the tree<br> c1:Number of events in the tree<br> c2:Average depth of leaves<br> c3:total number of leaves<br> c4:Sum of the depth of all leaves (=c2*c3)<br> c5:(=0) not used<br> c6:Magnitude of root<br> c7:Maximum magnitude of an event inside the tree</p> <p><br> 4.3. Tree-size frequencies<br> --------------------------</p> <p>18 x "*TopoTrees.FK" files obtained from "*TopoTrees.dat". Contains the frequencies of tree-sizes. Each raw number correspond to a size. Each value corresponds to number of incidences of that size divided by total number of events (10^5 in all cases). Notice that last point is missing at size = max-length, freq.= 1.0 / total number of events. Files used to generate fig. 3 of ref. [1]</p> <p><br> ################<br> 4. LIST OF FILES<br> ################</p> <p>(copy of this text)<br> readme.txt<br> md5:2073cdbb4a0afd8ab96f8bfaef20579f 13 Kb</p> <p>4.1. Raw data (42 files)<br> ------------------------</p> <p>ETASbranch_b0.50r0.00N0.50.Seq<br> md5:bae21678db4fe575b06b69a5f32253d2 12.5 Mb<br> ETASbranch_b0.50r0.00N0.99.Seq<br> md5:76ea70ba94aeeed7ba3c358305f4b85b 1.3 Gb<br> ETASbranch_b0.50r0.05N0.50.Seq<br> md5:0ccf5af07efb9dd99d63a9675fd3069f 12.4 Mb<br> ETASbranch_b0.50r0.05N0.99.Seq<br> md5:77a7bfd490af54a5a3d15c4183e9e6f7 1.4 Gb<br> ETASbranch_b0.50r0.10N0.50.Seq<br> md5:f274717039f8f69ce57d5054577777af 12.4 Mb<br> ETASbranch_b0.50r0.10N0.99.Seq<br> md5:e18f4c26acfdff8d79a6223aa86867d2 1.3 Gb<br> ETASbranch_b0.50r0.15N0.50.Seq<br> md5:5f93c14a8d0f88fd021bc6d8a5fc7d6e 12.3 Mb<br> ETASbranch_b0.50r0.15N0.99.Seq<br> md5:906e09ecb6c6c18fdbcfdca3d9dbe225 1.3 Gb<br> ETASbranch_b0.50r0.20N0.50.Seq<br> md5:2467c2c7fafeddaa630074e991cb7767 12.4 Mb<br> ETASbranch_b0.50r0.20N0.99.Seq<br> md5:64c45a47b7f953d7088b10d4badd068c 1.3 Gb<br> ETASbranch_b0.50r0.25N0.50.Seq<br> md5:d65abaa4664b2412707b27b6e9143215 12.4 Mb<br> ETASbranch_b0.50r0.25N0.99.Seq<br> md5:8d7467b0bd80ce7e8cbcaf5dfe2725b7 1.2 Gb<br> ETASbranch_b0.50r0.30N0.50.Seq<br> md5:fd3af9802b3b66a623172bb2aafc0a0b 12.4 Mb<br> ETASbranch_b0.50r0.30N0.99.Seq<br> md5:2ad1b9b9c858638067af6068dfd25d59 1.3 Gb<br> ETASbranch_b0.50r0.35N0.50.Seq<br> md5:8cf36cc90a19f1283c838dee4849626b 12.5 Mb<br> ETASbranch_b0.50r0.35N0.99.Seq<br> md5:6b35085100561865ac9463edfb166d7f 1.4 Gb<br> ETASbranch_b0.50r0.40N0.50.Seq<br> md5:317b154c8ed74262b18c41762597d236 12.3 Mb<br> ETASbranch_b0.50r0.40N0.99.Seq<br> md5:c83dc299035928c9ae8a7f6d101ac9b8 1.2 Gb<br> ETASbranch_b0.50r0.45N0.50.Seq<br> md5:d3baccfbf347e555c30f3dfb1db1d9c0 12.4 Mb<br> ETASbranch_b0.50r0.45N0.99.Seq<br> md5:5bf42163acc2eae1eeae8549b850ccc9 1.1 Gb<br> ETASbranch_b0.50r0.50N0.50.Seq<br> md5:3ad7427725dfcf302c2dc2c519dad8c0 12.4 Mb<br> ETASbranch_b0.50r0.50N0.99.Seq<br> md5:7316fb3661b87215a1a7b1c11a54975b 1.3 Gb<br> ETASbranch_b0.50r0.55N0.50.Seq<br> md5:6dbc9846506675c61d110ae918165a6b 12.4 Mb<br> ETASbranch_b0.50r0.55N0.99.Seq<br> md5:5d4d007bf898512185dd2f87b715715e 1.4 Gb<br> ETASbranch_b0.50r0.60N0.50.Seq<br> md5:569c1c409ef618f518fe6fb6f41493a3 12.7 Mb<br> ETASbranch_b0.50r0.60N0.99.Seq<br> md5:e739613598e8d2530d108ff24e8c045a 1.1 Gb<br> ETASbranch_b0.50r0.65N0.50.Seq<br> md5:e95cead475c1eedb0618bebf1b548ec2 12.3 Mb<br> ETASbranch_b0.50r0.65N0.99.Seq<br> md5:8fe8b042085f261971d94bf12f96dc7e 1 Gb<br> ETASbranch_b0.50r0.70N0.50.Seq<br> md5:48824f1efdbc31c5bb4d4f799bc166d5 12.2 Mb<br> ETASbranch_b0.50r0.70N0.99.Seq<br> md5:2ead803394a4c3b04f1e22e9b8c5ba46 872.9 Mb<br> ETASbranch_b0.50r0.75N0.50.Seq<br> md5:8c5cf4753836076c677a74d831354f33 12.2 Mb<br> ETASbranch_b0.50r0.75N0.99.Seq<br> md5:91ed931aedfb365761799e0882fcfb86 1 Gb<br> ETASbranch_b0.50r0.80N0.50.Seq<br> md5:f93e1779bc5344ffd17021bb387bb373 11 Mb<br> ETASbranch_b0.50r0.80N0.99.Seq<br> md5:6b0fd7fcad20ac4467fecda9502e9954 75.6 Mb<br> ETASbranch_b0.50r0.85N0.50.Seq<br> md5:a5debe21de70204a490c75dd7c68657a 11 Mb<br> ETASbranch_b0.50r0.85N0.99.Seq<br> md5:f9ed3ad6db12ad4a612c3945ccf72f18 48.5 Mb<br> ETASbranch_b0.50r0.90N0.50.Seq<br> md5:4493eef2170b68052cf56636c63091d7 9.7 Mb<br> ETASbranch_b0.50r0.90N0.99.Seq<br> md5:70043ed6c84e3896268c46f77eabefd3 26.7 Mb<br> ETASbranch_b0.50r0.95N0.50.Seq<br> md5:874916faa95c63312aa27aecea45de8f 7.5 Mb<br> ETASbranch_b0.50r0.95N0.99.Seq<br> md5:bbc54464c2e3cbac64401978216ae993 11.7 Mb<br> ETASbranch_b0.50r1.00N0.50.Seq<br> md5:3ff467b894f46a4be0b0747c686edad5 5.8 Mb<br> ETASbranch_b0.50r1.00N0.99.Seq<br> md5:d4d5416d56ed3de6eb2702a8eeb0a0b5 5.8 Mb</p> <p><br> 3.2. List of trees (53 files)<br> -----------------------------</p> <p><br> ETASbranch_b1.00r0.00N0.500TopoTrees.dat<br> md5:a5834b8d047fd3446e6d83488422bb5b 1.8 Gb<br> ETASbranch_b1.00r0.00N0.99TopoTrees.dat<br> md5:34eb899782444475151073ed6fcfc6aa 18.6 Mb<br> ETASbranch_b1.00r0.05N0.50TopoTrees.dat<br> md5:cc5f0edfc31a5f15394444243560e0ad 18.6 Mb<br> ETASbranch_b1.00r0.05N0.99TopoTrees.dat<br> md5:58cc7e7e746051f9efbe510c164a2b09 18.6 Mb<br> ETASbranch_b1.00r0.10N0.500TopoTrees.dat<br> md5:52cb3a9a5aeb1e222743bcdad3ea3b3d 1.8 Gb<br> ETASbranch_b1.00r0.10N0.50TopoTrees.dat<br> md5:a68a69280c0129ee3554d5c97dd0fa47 18.6 Mb<br> ETASbranch_b1.00r0.10N0.99TopoTrees.dat<br> md5:307edf1353af5141f11d89cbab6c4d30 18.6 Mb<br> ETASbranch_b1.00r0.15N0.500TopoTrees.dat<br> md5:52807004c32bdc1a999a2aaf1ff91bb9 1.8 Gb<br> ETASbranch_b1.00r0.15N0.50TopoTrees.dat<br> md5:0bdb45b13ed0fd5c0fa1619310d71670 18.6 Mb<br> ETASbranch_b1.00r0.15N0.99TopoTrees.dat<br> md5:f26e351184921ed9a85c8993404a4718 18.6 Mb<br> ETASbranch_b1.00r0.20N0.500TopoTrees.dat<br> md5:5fd930a98e28d551c9c204d22c6f564a 1.8 Gb<br> ETASbranch_b1.00r0.20N0.50TopoTrees.dat<br> md5:4aaeb952e0ecb7d4fc4bd2fb7822c729 18.6 Mb<br> ETASbranch_b1.00r0.20N0.99TopoTrees.dat<br> md5:a30687662b269f447d7e4cc020bd3773 18.6 Mb<br> ETASbranch_b1.00r0.25N0.500TopoTrees.dat<br> md5:4168e6e70b247977213d2278b94d65f3 1.8 Gb<br> ETASbranch_b1.00r0.25N0.50TopoTrees.dat<br> md5:b360877be00088b52676ba4717377761 18.6 Mb<br> ETASbranch_b1.00r0.25N0.99TopoTrees.dat<br> md5:34d11b64c78c852d4f50de7b1265c9b7 18.6 Mb<br> ETASbranch_b1.00r0.30N0.500TopoTrees.dat<br> md5:8c1a8b07c5f2350646d415a687f84492 1.8 Gb<br> ETASbranch_b1.00r0.30N0.50TopoTrees.dat<br> md5:0c1bfccd8cd141090a0bfb0cc7ae1ccd 18.6 Mb<br> ETASbranch_b1.00r0.30N0.99TopoTrees.dat<br> md5:652f975754241fed317acba066cf339b 18.6 Mb<br> ETASbranch_b1.00r0.35N0.500TopoTrees.dat<br> md5:6bc2a3f7b4f7fd8a92595271d2ea425f 1.8 Gb<br> ETASbranch_b1.00r0.35N0.50TopoTrees.dat<br> md5:d1d4e8a467d8d445cd3f6ca27c3c8af3 18.6 Mb<br> ETASbranch_b1.00r0.35N0.99TopoTrees.dat<br> md5:e56eee0119b193898ea729b75c572617 18.6 Mb<br> ETASbranch_b1.00r0.40N0.500TopoTrees.dat<br> md5:370a2f6c670fd7c9b515e43891417b73 1.8 Gb<br> ETASbranch_b1.00r0.40N0.50TopoTrees.dat<br> md5:9f50c23e33cca983df58e2b17f29f8e5 18.6 Mb<br> ETASbranch_b1.00r0.40N0.99TopoTrees.dat<br> md5:d9971930b46838de09a7ccdc7cd9b459 18.6 Mb<br> ETASbranch_b1.00r0.45N0.500TopoTrees.dat<br> md5:cfef99746043d0718ad19e48ffdfeee4 1.8 Gb<br> ETASbranch_b1.00r0.45N0.50TopoTrees.dat<br> md5:44667f56360024e40b11be4f52ccf9c3 18.6 Mb<br> ETASbranch_b1.00r0.45N0.99TopoTrees.dat<br> md5:09b84c6f3c9dd58331c6f9ab4eec8f58 18.6 Mb<br> ETASbranch_b1.00r0.50N0.500TopoTrees.dat<br> md5:e9e880f24ffc9ba40db7c7f78d0f8d89 1.8 Gb<br> ETASbranch_b1.00r0.50N0.50TopoTrees.dat<br> md5:ce4835d532513bffd54723f285207898 1.9 Mb<br> ETASbranch_b1.00r0.50N0.99TopoTrees.dat<br> md5:4c935d76aee113c88f4a3348197f75f3 18.6 Mb<br> ETASbranch_b1.00r0.55N0.500TopoTrees.dat<br> md5:46cd9e8d508c15009e90a194a141008f 1.8 Gb<br> ETASbranch_b1.00r0.55N0.50TopoTrees.dat<br> md5:8e1cb9c364c2a275d10373b54dab6239 18.6 Mb<br> ETASbranch_b1.00r0.55N0.99TopoTrees.dat<br> md5:304b947c52e826a9b4cb68c38f007443 18.6 Mb<br> ETASbranch_b1.00r0.60N0.500TopoTrees.dat<br> md5:df8b7889e76694188b29f7fc16f6dce2 1.8 Gb<br> ETASbranch_b1.00r0.60N0.50TopoTrees.dat<br> md5:fc6c2110e139290ae9401765e8aae782 18.6 Mb<br> ETASbranch_b1.00r0.60N0.99TopoTrees.dat<br> md5:6af466cc527b7422f1f958bfc6a87d15 18.6 Mb<br> ETASbranch_b1.00r0.65N0.500TopoTrees.dat<br> md5:129abd6de7821da465265a485217156d 1.8 Gb<br> ETASbranch_b1.00r0.65N0.50TopoTrees.dat<br> md5:469fc2ae3768fe9ae0ab0b8fbfaf3051 18.6 Mb<br> ETASbranch_b1.00r0.65N0.99TopoTrees.dat<br> md5:b7148c414b243b1911de8543d25e3d38 18.6 Mb<br> ETASbranch_b1.00r0.70N0.500TopoTrees.dat<br> md5:3a3eed181ae6308c29f274d5e14a97df 1.8 Gb<br> ETASbranch_b1.00r0.70N0.50TopoTrees.dat<br> md5:ca7472f17916e7f8634a97c08ca96c58 18.6 Mb<br> ETASbranch_b1.00r0.70N0.99TopoTrees.dat<br> md5:8c2a4ab3800e8a8178be5d6e856f5b50 18.6 Mb<br> ETASbranch_b1.00r0.75N0.50TopoTrees.dat<br> md5:374796f34a3e25f711601161f8a34ae7 18.6 Mb<br> ETASbranch_b1.00r0.75N0.99TopoTrees.dat<br> md5:591f6ac31545f3e7558b4a5e151c80a5 18.6 Mb<br> ETASbranch_b1.00r0.80N0.50TopoTrees.dat<br> md5:473c6e9ac33cc6e8eabe0ed3c91b40b7 18.6 Mb<br> ETASbranch_b1.00r0.80N0.99TopoTrees.dat<br> md5:81f41fbce89c931f87890133b4c0f9f9 18.6 Mb<br> ETASbranch_b1.00r0.85N0.50TopoTrees.dat<br> md5:2c38db4208898d95c1c5b2ef0a94c930 18.6 Mb<br> ETASbranch_b1.00r0.85N0.99TopoTrees.dat<br> md5:4866f179b52ff5a6191d47e6d8ead5ce 18.6 Mb<br> ETASbranch_b1.00r0.90N0.50TopoTrees.dat<br> md5:f14a7f700c0c73743eb5747399abdcce 18.6 Mb<br> ETASbranch_b1.00r0.90N0.99TopoTrees.dat<br> md5:da4bd00aeefcbd1e67f0d7fe8fb1d8be 18.6 Mb<br> ETASbranch_b1.00r0.95N0.50TopoTrees.dat<br> md5:57b7f0bd901bb47ba3c253f801301716 18.6 Mb<br> ETASbranch_b1.00r0.95N0.99TopoTrees.dat<br> md5:6381354768603c6d2129c99df2a7798a 18.6 Mb</p> <p>4.3. Tree-size frequencies (18 files)<br> -------------------------------------</p> <p>ETASbranch_b1.00r0.00N0.99TopoTrees.FK<br> md5:5b6ccaf1c1218f101ed24c332bfefec3 612 Kb<br> ETASbranch_b1.00r0.10N0.30TopoTrees.FK<br> md5:f9f07bbd7be5377e1589101e85640149 468 B<br> ETASbranch_b1.00r0.20N0.30TopoTrees.FK<br> md5:6c7e40f8b93c7a3e62f2c105f0e7a89b 558 B<br> ETASbranch_b1.00r0.20N0.99TopoTrees.FK<br> md5:158884c9bc4afa10c8ebb2e7b516a421 3.3 Mb<br> ETASbranch_b1.00r0.30N0.30TopoTrees.FK<br> md5:14d837b9763b9a54e64311e832e29f04 846 B<br> ETASbranch_b1.00r0.30N0.99TopoTrees.FK<br> md5:90b8b5a38a83cb4b4c833b605cf6b38e 2.1 Mb<br> ETASbranch_b1.00r0.40N0.30TopoTrees.FK<br> md5:1f198a7c08078066e7fa57ce9ebda8eb 3 Kb<br> ETASbranch_b1.00r0.40N0.99TopoTrees.FK<br> md5:229b837e7950fe85ea1c51be8e3f457e 3.3 Mb<br> ETASbranch_b1.00r0.50N0.30TopoTrees.FK<br> md5:6ead8bff09c3c2687c526648bfec6ab3 16 Kb<br> ETASbranch_b1.00r0.60N0.30TopoTrees.FK<br> md5:38c5cb9f544367ff9bdab766c4bcf03f 112 Kb<br> ETASbranch_b1.00r0.60N0.99TopoTrees.FK<br> md5:2591d4e2dd1c2051c0b7d96f423c526a 106.5 Mb<br> ETASbranch_b1.00r0.70N0.30TopoTrees.FK<br> md5:4f7741314e1a0d0b1c1733532d909277 7.5 Mb<br> ETASbranch_b1.00r0.80N0.30TopoTrees.FK<br> md5:3c4240e79ee19c03afac91d631f38cb8 602 Kb<br> ETASbranch_b1.00r0.80N0.30TopoTrees.FK<br> md5:3c4240e79ee19c03afac91d631f38cb8 602 Kb<br> ETASbranch_b1.00r0.90N0.30TopoTrees.FK<br> md5:914a782fa2e73c0bc5a6e70bbb2afec9 1.3 Mb<br> ETASbranch_b1.00r0.90N0.99TopoTrees.FK<br> md5:cc165045d4d9f25c0f9c30c0ec71759f 864 Kb<br> ETASbranch_b1.00r0.99N0.30TopoTrees.FK<br> md5:a08cfc27a0e42d54f207ea05b7210714 187 Kb<br> ETASbranch_b1.00r0.99N0.99TopoTrees.FK<br> md5:a15681be56f870a95fe62794b11524df 365 Kb</p> <p><br> 4.4 Known missing/broken files<br> ------------------------------</p> <p>ETASbranch_b1.00r0.50N0.50TopoTrees.dat<br> ETASbranch_b1.00r0.10N0.99TopoTrees.FK<br> ETASbranch_b1.00r0.20N0.99TopoTrees.FK<br> ETASbranch_b1.00r0.50N0.99TopoTrees.FK<br> ETASbranch_b1.00r0.70N0.99TopoTrees.FK<br> ETASbranch_b1.00r0.05N0.500TopoTrees.dat<br> ETASbranch_b1.00r0.70N0.500TopoTrees.dat<br> ETASbranch_b1.00r0.75N0.500TopoTrees.dat<br> ETASbranch_b1.00r0.80N0.500TopoTrees.dat<br> ETASbranch_b1.00r0.85N0.500TopoTrees.dat<br> ETASbranch_b1.00r0.90N0.500TopoTrees.dat<br> ETASbranch_b1.00r0.95N0.500TopoTrees.dat<br> </p>
A Bayesian Phylogenetic Hidden Markov Model for B Cell Receptor Sequence Analysis
<p>simulation and PC64/VRC01 input/output data files</p>
Models from: Exploring ensemble applications for multi-sequence myocardial pathology segmentation
<p>Trained models for the MyoPS2020 challenge http://www.sdspeople.fudan.edu.cn/zhuangxiahai/0/MyoPS20/</p> <p>As described in the publication: "Exploring ensemble applications for multi-sequence myocardial pathology segmentation"</p> <p>Source code available at: https://github.com/chfc-cmi/miccai2020-myops</p>
VFTS meeting: animation of main-sequence model evolution
<p>This animation shows the evolution of our binary and single stellar models from 2Myr to 100Myr. We populate 3763 binaries, whose primary mass is in the range of 3Msun to 100Msun, following a Salpeter IMF with an exponent of -2.37. The mass ratio is uniformly distributed from 0.1 to 1. The orbital period logP is uniformly distributed from the minimum value at which the two stars would contact initially to 3.5. Both components rotate at half of their critical velocities initially. </p> <p>Considering a binary fraction of 70%, we populate 1612 single stars with vi=0.5. In addition, we also populate 537 slowly-rotating single stars with vi=0.2 to reproduce the observed blue MS in young star clusters. The number is chosen such that the ratio between the slow and fast rotators is 1/3, which is the same as the ratio of observed blue and red MS stars. </p> <p>Gravity darkening and observational errors are included. Filled circles correspond to single stellar models, while open symbols correspond to binaries with different companions. Single stellar tracks with vi=0.5 are plotted with solid grey lines. The black dotted line represents the ZAMS line of vi=0.2 single stellar models. In the legend, the number in each group outside the parenthesis corresponds to the number of stars whose color and magnitude are in the figure range. While the number in the parenthesis corresponds to the number of stars above the orange dashed line, which is 1.75 mag below the turn-off magnitude. </p>
A codon model for associating phenotypic traits with altered selective patterns of sequence evolution
<p>Detecting the signature of selection in coding sequences and associating it with shifts in phenotypic states can unveil genes underlying complex traits. Of the various signatures of selection exhibited at the molecular level, changes in the pattern of selection at protein coding genes have been of main interest. To this end, phylogenetic branch-site codon models are routinely applied to detect changes in selective patterns along specific branches of the phylogeny. Many of these methods rely on a pre-specified partition of the phylogeny to branch categories, thus treating the course of trait evolution as fully resolved and assuming that phenotypic transitions have occurred only at speciation events. Here we present TraitRELAX, a new phylogenetic model that alleviates these strong assumptions by explicitly accounting for the uncertainty in the evolution of both trait and coding sequences. This joint statistical framework enables the detection of changes in selection intensity upon repeated trait transitions. We evaluated the performance of TraitRELAX using simulations and then applied it to two case studies. Using TraitRELAX, we found an intensification of selection in the primate SEMG2 gene in polygynandrous species compared to species of other mating forms, as well as changes in the intensity of purifying selection operating on sixteen bacterial genes upon transitioning from a free-living to an endosymbiotic lifestyle.</p>
Cross-disease integration of single-cell RNA sequencing data from lung myeloid cells reveals TAM signature in in vitro model
<p>Single cells from a 3D human cell-based model comprising tumor cell line-derived spheroids, cancer-associated fibroblasts and primary monocytes were dissociated and analyzed using scRNAseq. 4 monocyte donors were used in the 3D model, and 3 monocyte donors were used for 2D differentiation of macrophages.</p>
Automatic message sequence chart creation from simulation run of the proposed parametric colored Petri net model of the Chandy-Lamport algorithm
<p><span>These videos show the creation of two message sequence charts from simulation runs of the proposed parametric colored Petri net model of the Chandy-Lamport algorithm using the CPN tool with three constituting processes. </span></p> <p><strong><span>Message Sequence Chart of <span> </span>Parametric Model With 3 Processes via Automatic Simulation Run_SuppInfo.mp4</span></strong><span>: This video shows the automatic generation of a message sequence chart of the proposed parametric colored Petri net model of the Chandy-Lamport distributed global snapshot algorithm using the CPN tool version 4.0.0. The number of constituting processes is parametric in the model and was set to three. The video was generated using the authors' updated CPN tool extension server. The automatic simulation run of the model has been used to create this video. The CPN tool randomly selects the enabled transition at each step in an automatic simulation run.</span></p> <p><strong><span>Message Sequence Chart of Parametric Model With 3 Processes via Step-By-Step Simulation Run_SuppInfo.mp4:</span></strong><span> This video shows the automatic generation of a message sequence chart of the proposed parametric colored Petri net model of the Chandy-Lamport algorithm in a step-by-step simulation run with our updated extension server of the CPN tools version 4.0.0. The number of constituting processes is parametric in the model and was set to three. We manually fired our selected enabled transition of the model to create this video. </span></p>
Human breast cancer PDTX models bulk and single cell RNA sequencing
<p>This dataset includes information relevant to the following manuscript from the labs of Prof. Carlos Caldas (University of Cambridge), and Dr. Long V. Nguyen (Princess Margaret Cancer Centre, University Health Network):</p> <p>Nguyen LV et al. Dynamics and plasticity of human breast cancer single cell-derived clones. Under consideration for publication.</p> <p>Bulk RNA sequencing raw count matrices are provided (RawCounts.csv) along with the normalized count matrices (LogCPMNormCounts.csv).</p> <p>Single cell RNA sequencing count matrix processed from R package metacell is provided (mat.pdx_LN_v2_filt.Rda), along with the mc and mc2d files with information on metacell partitions (mc.pdx_LN_v2_filt.Rda and mc2d.pdx_LN_v2_filt.Rda).</p> <p>Single cell RNA sequencing count matrices processed using Seurat are also provided separately for each PDTX model analysed (STG139.rds, STG201.rds, AB040.rds and IC07.rds).</p> <p>Code and information on data analysis is provided for reviewers in our unpublished manuscript and on Github (https://github.com/cclab-brca/clone-dynamics).</p>
Single-Cell RNA-sequencing of neural precursor cells from an Alzheimer's mouse model, wild-type mice, and Alzheimer's mice rescued with Usp16 haploinsufficiency
<p class="MsoNormal">Alzheimer's disease (AD) is a progressive neurodegenerative disease observed with aging that represents the most common form of dementia. To date, therapies targeting end-stage disease plaques, tangles, or inflammation have limited efficacy. Therefore, we set out to identify an earlier targetable phenotype. Utilizing a mouse model of AD we found that cell intrinsic neural precursor cell (NPC) dysfunction precedes widespread inflammation and amyloid plaque pathology, making it one of the earlier defects in the evolution of the disease. We demonstrate that reversing impaired NPC self-renewal via genetic reduction of USP16, a histone modifier and critical physiological antagonist of the Polycomb Repressor Complex 1, can prevent downstream cognitive defects and decrease astrogliosis in vivo. To delineate potential self-renewal pathways that might contribute to the defect and rescue of Tg-SwDI NPCs and Tg-SwDI/<em>Usp16<sup><span>+/-</span></sup></em> NPCs, respectively, we performed single-cell RNA-seq and gene set enrichment analysis (GSEA) on lineage depleted primary FACS-sorted CD31<sup><span>-</span></sup>CD45<sup><span>-</span></sup>Ter119<sup><span>-</span></sup>CD24<sup><span>-</span></sup> NPCs from Tg-SwDI, WT, and Tg-SwDI/<em>Usp16<sup><span>+/-</span></sup></em> mice at 3-4 months and 1 year of age. Using the GSEA Hallmark gene sets, we found only three gene sets that were enriched in Tg-SwDI mice over WT mice and rescued in the Tg-SwDI/<em>Usp16<sup><span>+/-</span></sup> </em>mice at both ages: TGF-ß pathway, oxidative phosphorylation, and Myc Targets. The TGF-ß pathway consistently had the highest normalized enrichment score in pairwise comparisons between Tg-SwDI vs WT and Tg-SwDI vs Tg-SwDI/<em>Usp16<sup><span>+/-</span></sup> </em>of the three rescued pathways. These data suggest that USP16 may regulate neural precursor cell function in part through the BMP pathway.</p>
Sequence-dependent model of genes with dual σ factor preference
<p class="MsoNormal"><em>Escherichia coli</em> uses <span>s</span> factors to quickly control large gene cohorts during stress conditions. While most of its genes respond to a single <span>s</span> factor, approximately 5% of them have dual <span>s</span> factor preference. The most common are those responsive to both <span>s</span><sup>70</sup>, which controls housekeeping genes, and <span>s</span><sup>38</sup>, which activates genes during stationary growth and stresses. Using RNA-seq and flow-cytometry measurements, we show that 'σ<sup>70+38</sup> genes' are nearly as upregulated in stationary growth as 'σ<sup>38</sup> genes'. Moreover, we find a clear quantitative relationship between their promoter sequence and their response strength to changes in σ<sup>38</sup> levels. We then propose and validate a sequence dependent model of σ<sup>70+38</sup> genes, with dual sensitivity to <span>s</span><sup>38 </sup>and <span>s</span><sup>70</sup>, that is applicable in the exponential and stationary growth phases, as well in the transient period in between. We further propose a general model, applicable to other stresses and σ factor combinations. Given this, promoters controlling σ<sup>70+38</sup> genes (and variants) could become important building blocks of synthetic circuits with predictable, sequence-dependent sensitivity to transitions between the exponential and stationary growth phases.</p>
Consensus nucleotide sequences for env and gag for paper: Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models
<p>This is the consensus sequence repository to the manuscript "Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models".<br> It contains the 10% consensus nucleotide sequences of the env and gag (only p24) protein of HIV-1 used for the training and leftout data set. The NGS sequences are available under BioProject ID PRJNA810303 and the corresponding BioSample Accession IDs are SAMN26241863:26242168 and SAMN28728524:SAMN28728529</p> <ul> <li>env_leftout.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the env protein for the leftout data set</li> </ul> </li> <li>env_nt_274.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the env protein for the training data set</li> </ul> </li> <li>gag_leftout.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the gag protein for the leftout data set</li> </ul> </li> <li>gag_nt_274.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the gag protein for the training data set</li> </ul> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.