Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
59
datasets available to search
ShareScore release 0.9.0
Dataset results
59 results for “Syntax”
On-the-Fly Syntax Highlighting Using Neural Networks - Replication Package (Data)
<p>This dataset includes the data to replicate the study for the paper <em>On-the-Fly Syntax Highlighting Using Neural Networks</em>. It can be reused for future research in the field. We also include the detailed results obtained by executing our approach.</p> <p>HLNN-Resources.zip includes the input data already formatted to be directly used with the shared source code.</p> <p>The paper is published in the proceeding of the <em>30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE)</em>.</p>
Supplementary Materials for 'Measuring and assessing indeterminacy and variation in the morphology-syntax distinction'
<p><strong>Supplementary materials for the article 'Measuring and assessing indeterminacy and variation in the morphology-syntax distinction' in <em>Linguistic Typology </em>(Vol. and No. TBD).</strong></p> <p>Abstract:</p> <p>We provide a discussion of some of the challenges in using statistical methods to investigate the morphology-syntax distinction cross-linguistically. The paper is structured around three problems related to the morphology-syntax distinction; (i) the boundary strength problem; (ii) the composition problem; (iii) the architectural problem.<br> The boundary strength problem refers to the possibility that languages vary in terms of how distinct morphology and syntax are or the degree to which morphology is autonomous. The composition problem refers to the possibility that languages vary in terms of how they distinguish morphology and syntax: what types of properties distinguish the two systems. The architecture problem refers to the possibility that languages vary in terms of whether a global distinction between morphology and syntax is motivated at all and the possibility that languages might partition phenomena in different ways.<br> This paper is concerned with providing an overarching review of the methodological problems involved in addressing these three issues. We illustrate the problems using three statistical methods: correlation matrices, random forests with different choices for the dependent variable, and hierarchical clustering with validation techniques.</p> <p> </p> <p>Overview of materials:</p> <ul> <li>SM1: csv with the data</li> <li>SM2: code and pdf for generating the correlation matrices</li> <li>SM3: code and pdf for the random forest analyses</li> <li>SM4: code and pdf for the clustering and cluster validation analyses</li> </ul>
Neural mechanisms of musical syntax and tonality, and the effect of musicianship
Open the record for dataset details and reuse information.
Between syntax and morphology: German noun+verb units (Glossa)
<p><strong>This dataset accompanies a paper to be published in Glossa. Under the present DOI, all data generated for this research as well as all scripts used are stored. The paper itself is CC-licensed, refer to glossa-journal.org.</strong></p><p><strong>Abstract</strong></p><p>We show that graphemic variation—at least in some writing systems—can be analysed in terms of grammatical variation given a usage-based probabilistic view of the grammar-graphemics interface. Concretely, we examine a type of noun+verb unit in German, which can be written as one word or two. We argue that the variation in writing is rooted in the units' ambiguous status in between morphology (one word) and syntax (two words). The major influencing factors are shown to be the semantic relation between the noun and the verb (argument or oblique relation) and the morphosyntactic context. In prototypically nominal contexts, a re-interpretation of the unit as a noun+noun compound is facilitated, which favours spelling as one word, while in prototypically verbal contexts, a syntactic realisation and consequently spelling as two words is preferred. We report the results of two large-scale corpus studies and a controlled production experiment to corroborate our analysis.</p>
Reproduction Package (VirtualBox Image) for the POPL 2024 Article `Enhanced Enumeration Techniques for Syntax-Guided Synthesis of Bit-Vector Manipulations`
<p>This is the artifact for the ACM PACMPL article <i>Enhanced Enumeration Techniques for Syntax-Guided Synthesis of Bit-Vector Manipulations</i>. We provide our artifact as an easy-to-use VirtualBox image, which contains the benchmarks, our tools for bit-vector synthesis, and the scripts for generating the results showcased in the paper.</p>
On-the-Fly Syntax Highlighting: Generalisation and Speed-ups - Replication Package
<p><strong>On-the-Fly Syntax Highlighting: Generalisation and Speed-ups</strong></p> <p>On-the-fly syntax highlighting involves the rapid association of visual secondary notation with each character of a language derivation. This task has grown in importance due to the widespread use of online software development tools, which frequently display source code and heavily rely on efficient syntax highlighting mechanisms. In this context, resolvers must address three key demands: speed, accuracy, and development costs. Speed constraints are crucial for ensuring usability, providing responsive feedback for end users and minimizing system overhead. At the same time, precise syntax highlighting is essential for improving code comprehension. Achieving such accuracy, however, requires the ability to perform grammatical analysis, even in cases of varying correctness. Additionally, the development costs associated with supporting multiple programming languages pose a significant challenge. The technical challenges in balancing these three aspects explain why developers today experience significantly worse code syntax highlighting online compared to what they have locally. The current state-of-the-art relies on leveraging programming languages' original lexers and parsers to generate syntax highlighting oracles, which are used to train base Recurrent Neural Network models. However, questions of generalisation remain. This paper addresses this gap by extending previous work validation dataset to six mainstream programming languages thus providing a more thorough evaluation. In response to limitations related to evaluation performance and training costs, this work introduces a novel Convolutional Neural Network (CNN) based model, specifically designed to mitigate these issues. Furthermore, this work addresses an area previously unexplored performance gains when deploying such models on GPUs. The evaluation demonstrates that the new CNN-based implementation is significantly faster than existing state-of-the-art methods, while still delivering the same near-perfect accuracy.</p>
Western Thrace Turkish: Phonology - Phonetic Features, Morphology and Syntax
<p>These snippets present linguistic features of Western Thrace Turkish, a Balkan Turkish dialect spoken in Northeastern Greece. This lecture is part of the lecture series: <em>Glottothèque: Languages of the Anatolia, Caucasus, Iran, Mesopotamia; grammatical snippets online </em>(electronic resource). Bamberg, Cambridge, Göttingen, Moskow, Nicosia, Paris: LACIM network, at https://spw.uni-goettingen.de/projects/lacim/, edited by Christiane Bulut, Anaïd Donabédian-Demopoulos, Geoffrey Haig, Geoffrey Khan, Pollet Samvelian, Stavros Skopeteas, Nina Sumbatova.</p>
WCRP baseline variable syntax
<p>The syntax used to describe the baseline variables presented in Juckes et al., 2023 (in preparation).</p> <ul> <li>Time sampling title: a brief description of the temporal sampling.</li> <li>Spatial sampling title: a brief description of the spatial sampling.</li> <li>Spatial sampling label: a label associated with the spatial dimensions of the variable.</li> <li>Additional Coordinates [optional]: listing of additional spatial and masking coordinates,</li> <li>Time sampling label: a label associated with the temporal dimensions and sampling of the variable.</li> <li>Comment [optional]: additional clarification.</li> </ul>
Data from: Both learning and syntax recognition are used by great tits when answering to mobbing calls
<p><span>Mobbing behavior, in addition to its complex cooperative aspects, is particularly suitable to study the mechanisms implicated in heterospecific communication. Indeed, various mechanisms ranging from pure learning to innate recognition have been proposed. One promising, yet understudied mechanism could be syntax recognition, especially given the latest works published on syntax comprehension in birds. In this experiment, we test whether great tits use both learning and syntax recognition when responding to heterospecifics. In the first part of the experiment, we demonstrate that great tits show different responses to the same heterospecific calls depending on their sympatric status. In a second part, we explore the impact of reorganizing the notes of the heterospecific mobbing calls to fit the syntax of great tits. Great tits showed an increased mobbing response toward the heterospecific calls when they shared their own call organization. Our results corroborate the recent finding that syntactic rules in bird calls may have a strong impact on their communication systems and enlighten how various mechanisms can be used by the same species to respond to heterospecific calls.</span></p>
Database and Syntax for Analysis of the Paper: "Effects of introducing the WHO Labour Care Guide on Caesarean section: a pragmatic, stepped-wedge, cluster randomized trial in India"
<p>The following files contains the information used to analyze the trial “Implementing the WHO Labour Care Guide to reduce the use of Caesarean section in four hospitals in India: a pragmatic, stepped wedge, cluster randomized pilot trial” in which it was hypothesized that the intervention would promote correct LCG use by these providers, changing their labour monitoring and management practices to align with WHO’s intrapartum recommendations. In turn, this could reduce overuse of Caesarean section, improve maternal and newborn outcomes, and enhance women’s care experiences. </p> <p>Two datafiles with extension “csv” are uploaded. The databased named “LCG Trial Women Database (transition period included).csv” is the database which contains the data of the recruited women in the trial. There is one row per women. The databased named “LCG Trial Neonates Database (transition period included).csv” is the database which contains the data of the neonates born from the recruited women. There is one row per neonate.</p> <p>The excel file “Data Dictionary LCG to Share.xlsx” is the data dictionary of the two databases. In the sheet named “Maternal Variables” a list and description of the variables included in the maternal database is included and, in the sheet, named “Neonatal Variables” a list and description of the variables included in the neonatal database is included.</p> <p>Three files of “R” extension and one “rmd” are included. The file named “RunningModelsFunctions.R” is the one use to run the models that are included in the analyses, the file named “2. Final Analysis LCG Trial.R” is the one in which the tables are prepared, and the file named “3. LCG Results Output Final.rmd” is used to export the tables with results. The R file named “funciones.tablas.R” is used in the analyses.</p>
Analysis Products: Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency
<p>This record contains analysis products for the paper "Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency" by Nair, Ameen <em>et al</em>. Please refer to the READMEs in the directories, which are summarized below.</p> <p>The record contains the following files:<br> <br> `clusters.tsv`: <strong> </strong>contains the cluster id, name and colour of clusters in the paper</p> <p><strong>scATAC.zip</strong></p> <p>Analysis products for the single-cell ATAC-seq data. Contains:</p> <p>- `cells.tsv`: list of barcodes that pass QC. Columns include:<br> - `barcode`<br> - `sample`: (time point)<br> - `umap1`<br> - `umap2`<br> - `cluster`<br> - `dpt_pseudotime_fibr_root`: pseudotime values treating a fibroblast cell as root<br> - `dpt_pseudotime_xOSK_root`: pseudotime values treating xOSK cell as root<br> - `peaks.bed`: list of peaks of 500bp across all cell states. 4th column contains the peak set label. Note that ~5000 peaks are not assigned to any peak set and are marked as NA.<br> - `features.tsv`: 50 dimensional representation of each cell <br> - `cell_x_peak.mtx.gz`: sparse matrix of fragment counts within peaks. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (combine sample + barcode). Rows correspond to peaks in `peaks.bed` </p> <p><strong>scATAC_clusters.zip</strong></p> <p>Analysis products corresponding to cluster pseudo-bulks of the single-cell ATAC-seq data. </p> <p>- `clusters.tsv`: contains the cluster id, name and colour used in the paper<br> - `peaks`: contains `overlap_reproducibilty/overlap.optimal_peak` peaks called using ENCODE bulk ATAC-seq pipeline in the narrowPeak format.<br> - `fragments`: contains per cluster fragment files </p> <p><strong>scATAC_scRNA_integration.zip</strong></p> <p>Analysis products from the integration of scATAC with scRNA. Contains:</p> <p>- `peak_gene_links_fdr1e-4.tsv`: file with peak gene links passing FDR 1e-4. For analyses in the paper, we filter to peaks with absolute correlation >0.45.<br> - `harmony.cca.30.feat.tsv`: 30 dimensional co-embedding for scATAC and scRNA cells obtained by CCA followed by applying Harmony over assay type.<br> - `harmony.cca.metadata.tsv`: UMAP coordinates for scATAC and scRNA cells derived from the Harmony CCA embedding. First column contains barcode.</p> <p><strong>scRNA.zip</strong></p> <p>Analysis products for the single-cell RNA-seq data. Contains:</p> <p>- `seurat.rds`: seurat object that contains expression data (raw counts, normalized, and scaled), reductions (umap, pca), knn graphs, all associated metadata. Note that barcode suffix (1-9 corresponds to samples D0, D2, ..., D14, iPSC)<br> - `genes.txt`: list of all genes<br> - `cells.tsv`: list of barcodes that pass QC across samples. Contains:<br> - `barcode_sample`: barcode with index of sample (1-9 corresponding to D0, D2, ..., D14, iPSC) <br> - `sample`: sample name (D0, D2, .., D14, iPSC)<br> - `umap1`<br> - `umap2`<br> - `nCount_RNA`<br> - `nFeature_RNA`<br> - `cluster`<br> - `percent.mt`: percent of mitochondrial transcripts in cell<br> - `percent.oskm`: percent of OSKM transcripts in cell<br> - `gene_x_cell.mtx.gz`: sparse matrix of gene counts. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (barcode suffix contains sample information). Rows correspond to genes in `genes.txt` <br> - `pca.tsv`: first 50 PC of each cell<br> - `oskm_endo_sendai.tsv`: estimated raw counts (cts, may not be integers) and log(1+ tp10k) normalized expression (norm) for endogenous and exogenous (Sendai derived) counts of POU5F1 (OCT4), SOX2, KLF4 and MYC genes. Rows are consistent with `seurat.rds` and `cells.tsv`</p> <p><strong>multiome.zip</strong></p> <p><em>multiome/snATAC:</em></p> <p>These files are derived from the integration of nuclei from multiome (D1M and D2M), with cells from day 2 of scATAC-seq (labeled D2). </p> <p>- `cells.tsv`: This is the list of nuclei barcodes that pass QC from multiome AND also cell barcodes from D2 of scATAC-seq. Includes:<br> - `barcode`<br> - `umap1`: These are the coordinates used for the figures involving multiome in the paper.<br> - `umap2`: ^^^ <br> - `sample`: D1M and D2M correspond to multiome, D2 corresponds to day 2 of scATAC-seq<br> - `cluster`: For multiome barcodes, these are labels transfered from scATAC-seq. For D2 scATAC-seq, it is the original cluster labels. <br> - `peaks.bed`: This is the same file as scATAC/peaks.bed. List of peaks of 500bp. 4th column contains the peak set label. Note that ~5000 peaks are not assigned to any peak set and are marked as NA.<br> - `cell_x_peak.mtx.gz`: sparse matrix of fragment counts within peaks. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (combine sample + barcode). Rows correspond to peaks in `peaks.bed`.<br> - `features.no.harmony.50d.tsv`: 50 dimensional representation of each cell prior to running Harmony (to correct for batch effect between D2 scATAC and D1M,D2M snMultiome). Rows correspond to cells from `cells.tsv`.<br> - `features.harmony.10d.tsv`: 10 dimensional representation of each cell after running Harmony. Rows correspond to cells from `cells.tsv`.</p> <p><em>multiome/snRNA:</em></p> <p>- `seurat.rds`: seurat object that contains expression data (raw counts, normalized, and scaled), reductions (umap, pca),associated metadata. Note that barcode suffix (1,2 corresponds to samples D1M, D2M). Please use the UMAP/features from snATAC/ for consistency.<br> - `genes.txt`: list of all genes (this is different from the list in scRNA analysis)<br> - `cells.tsv`: list of barcodes that pass QC across samples. Contains:<br> - `barcode_sample`: barcode with index of sample (1,2 corresponding to D1M, D2M respectively) <br> - `sample`: sample name (D1M, D2M)<br> - `nCount_RNA`<br> - `nFeature_RNA`<br> - `percent.oskm`: percent of OSKM genes in cell<br> - `gene_x_cell.mtx.gz`: sparse matrix of gene counts. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (barcode suffix contains sample information). Rows correspond to genes in `genes.txt` </p>
SYNTAX III REVOLUTION Trial: A Randomized Study Investigating the Use of CT Scan and Angiography of the Heart to Help the Doctors Decide Which Method is the Best to Improve Blood Supply to the Heart i
ClinicalTrials.gov study NCT02813473. IPD Sharing: YES. Countries: 5. Publications: 2.
Female audience shapes the complexity and syntax of male courtship displays in a lek-mating bird
Open the record for dataset details and reuse information.
Data from: Both learning and syntax recognition are used by great tits when answering to mobbing calls
Open the record for dataset details and reuse information.
Identifying Energy Efficiency Patterns in Sorting Algorithms via Abstract Syntax Tree Mining
<p>Replication package for Submission "Identifying Energy Efficiency Patterns in Sorting Algorithms via Abstract Syntax Tree Mining".</p> <p>Authors kept anonymous for review.</p>
Dissection of core promoter syntax through single nucleotide resolution modeling of transcription initiation (CLIPNET data)
<div>This contains data necessary to reproduce the figures in the CLIPNET paper (preprint <a href="https://www.biorxiv.org/content/10.1101/2024.03.13.583868">here</a>) as well as processed data used to train and evaluate CLIPNET. To preserve subdirectory structure, we've packaged the data into tar archives. Please refer to the README documents in our manuscript GitHub repo for more details on file contents: <a href="https://github.com/Danko-Lab/clipnet_paper/">https://github.com/Danko-Lab/clipnet_paper/</a></div> <div> </div> <div>Pretrained CLIPNET models are archived separately at <a href="../doi/10.5281/zenodo.10408622">DOI 10.5281/zenodo.10408622</a></div> <div> </div> <div>V5: Fixed bug in calculation of profile attribution scores causing them to be off by a factor of exactly 500. Genome-wide DeepSHAP tracks & TF-MoDISco tracks have been accordingly updated. I have not updated the individual examples, as these can be quickly fixed by simply multiplying by 500 when plotting. Additionally, I have uploaded profile and quantity motif calls, which contain genome-wide seqlet annotations. The columns in these files are [chrom, start, end, peak_idx, motif_annotation].</div> <div>V4: Uploaded individual bigWigs. These have been lifted over using CrossMap from the original hg19 (GSE110638) to hg38 and RPM normalized.</div> <div>V3: Final version prior to journal submission. Don't recall exact details of what's changed.</div> <div>V2: evaluation_metrics.tar.gz and evaluation_data.tar.gz have been replaced. Previously, we benchmarked the models by treating each peak in each individual as a separate data point. Here, we instead predicted from the reference genome and compared against the averaged bigWigs.</div>
Dataset, SPSS syntax, and SEM model for study 2 paper Implicit Agentic Narcissism
Open the record for dataset details and reuse information.
Long-distance dependencies in birdsong syntax
<p><span><span><span><span><span><span><span><span><span><span><span>Songbird syntax is generally thought to be simple, lacking in particular long-distance dependencies in which one element affects choice of another occurring considerably later in the sequence. Here we test for long-distance dependencies in the sequences of songs produced by song sparrows (<i>Melospiza melodia</i>). Song sparrows sing with eventual variety, repeating each song type in a consecutive series termed a "bout." We show that in switching between song types, song sparrows follow a "cycling rule," cycling through their repertoires in close to the minimum possible number of bouts. Song sparrows do not cycle in a set order but rather vary the order of song types from cycle to cycle. Cycling in a variable order strongly implies long-distance dependencies: choice of the next type must depend on the song types sung over the past cycle, in the range of 9-10 bouts. Song sparrows also follow a "bout length rule," whereby the number of repetitions of a song type in a bout is positively associated with the length of the interval until that type recurs. This rule requires even longer-distance dependencies that cross one another; such dependencies are characteristic of more complex levels of syntax than previously attributed to non-human animals.</span></span></span></span></span></span></span></span></span></span></span></p>
Mapped data: Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency
<p>This record contains mapped sequencing data for the paper "Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency" by Nair, Ameen <em>et al</em>. It contains single-cell RNA-seq (scRNA) and single-cell ATAC-seq (scATAC) data from a time course of human dermal fibroblasts induced with Yamanaka factors OSKM using a Sendai virus based delivery system. The scRNA and scATAC data is performed at days 0, 2, 4, 6, 8, 10, 12, 14 and the final iPSCs. The experiment was re-performed and single-nucleus multiome (ATAC+RNA) was collected on days 1 and 2. </p> <p>The data is as follows:</p> <p><strong>scATAC</strong>: We used Chromap (commit <a href="https://github.com/haowenz/chromap/tree/6e97125b9">https://github.com/haowenz/chromap/tree/6e97125b9</a>, <a href="https://doi.org/10.1038/s41467-021-26865-w">https://doi.org/10.1038/s41467-021-26865-w</a>) to perform barcode correction, alignment and filtering for each of our samples. The corresponding fragment files (tab separated file containing mapped fragments with columns: chr, start, end, barcode, number of reads) and their tabix indices are available for each sample.</p> <p><strong>scRNA</strong>: We used cellranger v6.0.2 for read mapping and quantification to obtain the counts matrix. We used the GRCh38 2020-A reference. For each sample, the raw and filtered counts matrices are provided. E.g. `D0/raw_feature_bc_matrix.h5` contains an HDF5 object containing gene counts for each barcode and associated metadata for the Day 0 sample. Similarly, the files in `D0/raw_feature_bc_matrix/` contain the same gene x barcode matrix, with the counts matrix in Matrix Market format (`matrix.mtx.gz`), and gene (`features.tsv.gz`) and barcode names (`barcodes.tsv.gz`). </p> <p><strong>multiome</strong>: The ATAC and RNA components are separately processed using the same tools as mentioned above for scATAC and scRNA. Outputs are in the `snATAC` and `snRNA` subdirectories respectively. In addition, the `ATAC.RNA.bc.map.tsv` file contains a map to link snATAC barcodes to snRNA barcodes. </p>
ChromBPNet models and data: Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency
<p>This record contains ChromBPNet models and data used to train the models for the paper "Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency" by Nair, Ameen <em>et al</em>.</p> <p>`data` contains bigwigs and regions (peaks + non-peaks) used for training each of the models. See `data/README.txt` for more details.</p> <p><strong>Models:</strong></p> <p><em>Loading the model:</em></p> <p>The models were trained using tf1.14. The models are provided in h5 format for tf1.14 (py3.7) and SavedModel format for tf2.X. tf2.X tested only for py3.8-11, tf2.8-13.</p> <p>To load the models in tf1.14:</p> <pre><code class="language-python">model = tf.keras.models.load_model("path/to/model.h5")</code></pre> <p>In tf2:</p> <pre><code class="language-python">model = tf.keras.models.load_model("path/to/model_dir")</code></pre> <p>If all fails, you can load the architecture as provided in `model_arch.py` with default parameters (`bpnet_seq` for bias model and `chrombpnet` for chrombpnet model), and then load the weights using `model.load_weights` from the weights provided in the `weights` directory.</p> <p> </p> <p><em>Usage:</em></p> <p>The bias models take as input one-hot sequence of length 2000. It has 2 outputs, a vector of logits of length 2000, and 1 logcounts scalar:</p> <pre><code class="language-python"># seq_one_hot of length B x 2000 x 4 out_bias_logits, out_bias_logcounts = bias_model.predict(seq_one_hot) # out_bias_logits: B x 2000 # out_bias_logcounts: B x 1</code></pre> <p>The ChromBPNet model takes as input a one-hot sequence of length 2000, bias logits of length 2000 and bias log-counts scalar. It has the same output types as the bias model. To run the chrombpnet model to obtain predictions:</p> <pre><code class="language-python">pred_profile, pred_logcounts = chrombpnet_model.predict([seq_one_hot, out_bias_logits, out_bias_logcounts]) # pred_profile: B x 2000 # pred_logcounts: B x 1 </code></pre> <p>If you wish to obtain the "de-biased" predictions (see Methods), simply pass in zeros instead of the bias model predictions as:</p> <pre><code class="language-python">pred_profile_debiased, pred_logcounts_debiased = chrombpnet_model.predict([seq_one_hot, np.zeros((seq_one_hot.shape[0], 2000)), np.zeros((seq_one_hot.shape[0], 1))])</code></pre> <p>To obtain predicted per-base predicted counts (with or without bias):</p> <pre><code class="language-python">pred_per_base_counts = scipy.special.softmax(pred_profile, axis=-1) * (np.exp(pred_logcounts)-1) # pred_per_base_counts: B x 2000 </code></pre> <p>Note that in general predicted counts can't be compared across models as they are not corrected for sequencing depth.</p> <p> </p> <p><em>Note:</em></p> <p>All bias models used across folds are identical, except for the final intercept term in the counts output (see Methods), that is specific to each cell state, fold combination.</p> <p> </p> <p><em>Folds:</em></p> <p>The splits used for training the different folds are as below:</p> Fold Test Chromosomes Validation Chromosomes 0 chr1 chr8, chr10 1 chr2, chr19 chr1 2 chr3, chr20 chr2, chr19 3 chr6, chr13, chr22 chr3, chr20 4 chr5, chr16, chrY chr6, chr13, chr22 5 chr4, chr15, chr21 chr5, chr16, chrY 6 chr7, chr18, chr14 chr4, chr15, chr21 7 chr11, chr17, chrX chr7, chr18, chr14 8 chr9, chr12 chr11, chr17, chrX 9 chr8, chr10 chr9, chr12 <p>Remaining chromosomes were used as the training chromosome for each fold.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.