Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,199
datasets available to search
ShareScore release 0.7.1
Dataset results
1,199 results for “alignment”
Coat protein (CP) and trimmed replication-associated protein (Rep) amino acid alignments, phylogenetic analyses, and associated metadata for ICTV-approved begomovirus RefSeq species exemplars
<p>DATA RETRIEVAL</p> <p>Annotated begomovirus coding sequences corresponding to each begomovirus species exemplar with a RefSeq accession number listed in the ICTV Virus Metadata Resource (VMR #18, 2021-10-19, <a href="https://ictv.global/vmr">https://ictv.global/vmr</a>) were downloaded from GenBank in protein FASTA file format. CP and Rep amino acid sequences were extracted and split into separate data sets for analysis. We confirmed the identity of misannotated ORF products by performing a BLAST search. For exemplar sequences missing ORF annotations (listed in metadata spreadsheet), ORFfinder (<a href="https://www.ncbi.nlm.nih.gov/orffinder/">https://www.ncbi.nlm.nih.gov/orffinder/</a>) was used to identify CP and Rep ORFs that were subsequently translated and added to each corresponding data set after BLAST confirmation.</p> <p>ALIGNMENTS</p> <p>Multiple sequence alignments were constructed using the MUSCLE method (Edgar, 2004) as implemented in MEGA 11 (Tamura et al., 2021) and manually corrected using AliView v1.26<strong> </strong>(Larsson, 2014). After an initial alignment inspection, exemplars with either severely truncated (i.e., length < 50% of the average length of the protein) or very divergent (i.e., causing us to doubt protein homology) CP or Rep sequences were excluded from the data set. Due to the difficulties in aligning the Rep sequences at the N- and C- terminal ends, the Rep alignment was trimmed to eliminate all residues prior to the iteron related domain (i.e., the known Rep functional region closest to the Rep start (Arguello-Astorga & Ruiz-Medrano, 2001)) in the N-terminus and after a conserved geminivirus motif found near the C-terminus, which corresponds to where other circular, Rep-encoding single-stranded DNA viruses possess an arginine finger motif (Kazlauskas et al., 2019; Krupovic et al., 2020). In total, our CP and Rep data sets contained amino acid sequences from 432 begomovirus species exemplars that met our inclusion criteria.</p> <p>PHYLOGENETIC ANALYSIS</p> <p>Maximum likelihood (ML) trees were inferred with IQ-Tree v2.0.7 (Minh et al., 2020) using the best fitting substitution model identified by the built-in ModelFinder feature (Kalyaanamoorthy et al., 2017). Tree inference was performed with 3000 ultrafast bootstrap (UFBoot) replicates, a perturbation strength of 0.2 and a stopping rule requiring an iteration interval of 500 iterations between unsuccessful improvements to the local optimum. The -bnni flag was enabled to reduce the risk of overestimating branch supports with UFBoot due to severe model violations. The provided phylogenies in NEXUS format are midpoint-rooted and branches are colored based on traditional begomovirus geographic groupings: exemplars sampled in the Americas in orange and exemplars sampled in the 'Africa, Asia, Europe and Oceania' (AAEO) region in blue. </p> <p>METADATA</p> <p>Metadata associated with each ICTV-approved species exemplar (n=445) – including country of isolation, geographic designation (i.e., AAEO/Americas), genome segmentation (i.e., monopartite/bipartite), presence/absence of V2/AV2 gene and length of genome/DNA-A segments – are included. Exemplars not incorporated into the other analyses are highlighted in red on the spreadsheet.</p> <p> </p>
CLDF dataset derived from List and Prokić's "Benchmark Database of Phonetic Alignments" from 2014
<p>Cite the source of the dataset as:</p> <blockquote> <p>List, Johann-Mattis and Jelena Prokić. (2014). A benchmark database of phonetic alignments in historical linguistics and dialectology. In: Proceedings of the International Conference on Language Resources and Evaluation (LREC), 26 — 31 May 2014, Reykjavik. 288-294.</p> </blockquote>
Translation Alignment: Ancient Greek to English. Annotation Style Guide and Gold Standard.
<p>This dataset contains guidelines and a gold standard for the alignment of Ancient Greek texts with English translations.</p> <p>The guidelines were used to annotate a diverse dataset including Homeric epic, Attic prose, and Platonic dialogue, and were tested by measuring inter-annotator agreement of 80% or higher. The Ancient Greek texts used are almost entirely available through the Scaife viewer (<a href="https://scaife.perseus.org/">https://scaife.perseus.org/</a>).</p> <p>The datasets used to develop the gold standard were aligned using the Ugarit Translation Alignment Editor for Historical languages (<a href="http://ugarit.ialigner.com/">http://ugarit.ialigner.com/</a>).</p> <p>The materials available here can be used to perform and evaluate alignments of various texts in Ancient Greek, to create new gold standard corpora, and to train automated translation models.</p> <p>The guidelines can also be further adapted to address similar language pairs including an inflected and a synthetic language, such as Latin and English, or can provide a structure for the alignment of other historical texts against modern translations. However, the guidelines are not project-specific: they were specifically intended for the creation of a Gold Standard in the scenario of machine translation. Different scenarios, such as language research or pedagogy, may need further tweaking to these guidelines to make them more compatible with different underlying principles.</p> <p>For further information on Ugarit and translation alignment of historical languages, see <a href="http://ugarit.ialigner.com/bib.php">http://ugarit.ialigner.com/bib.php</a> and follow us on Twitter (@ugarit_ty).</p>
Microwave Single Scattering Properties Database (Horizontally Aligned Aggregates of Dendrites)
<p>The database contains physical and microwave single scattering properties of horizontally aligned frozen hydrometeors as large as 11 cm in diameter. </p> <p>A description of the aggregation model used for particle generation can be found in:<br> Leinonen, J., and Szyrmer, W. (2015), Radar signatures of snowflake riming: A modeling study, <em>Earth and Space Science</em>, 2, 346– 358, doi:<a href="https://doi.org/10.1002/2015EA000102">10.1002/2015EA000102</a>.<br> The code used for particle generation is freely available at: <a href="https://github.com/jleinonen/aggregation">https://github.com/jleinonen/aggregation</a></p> <p>The scattering properties of particles were computed using discrete dipole approximation using ADDA software package (<a href="https://github.com/adda-team/adda">https://github.com/adda-team/adda</a>)</p> <p>Terminal velocity of snowflakes was computed using 4 hydrodynamical models that were implemented as a part of snowScat library (<a href="https://github.com/OPTIMICe-team/snowScatt">https://github.com/OPTIMICe-team/snowScatt</a>)</p> <p>Approximately one half of the snowflake structure files and one quarter of scattering properties (for X, Ku, Ka and W band) were generated for the publication of Leinonen and Szyrmer (2015). The remaining part of the dataset was generated using the ALICE High Performance Computing Facility at the University of Leicester.</p>
Alignments from "Caecilian genomes reveal molecular basis of adaptation and convergent evolution of limblessness in vertebrates"
<p>Compressed file containing the alignments at both nucleotide and amino acid level for the manuscript "Caecilian genomes reveal molecular basis of adaptation and convergent evolution of limblessness in vertebrates" </p>
Multiple Sequence Alignment of a diverse dataset with 1788 Mycobacterium tuberculosis isolates
<p><strong>Multiple Sequence Alignment of a diverse dataset with 1788 <em>Mycobacterium tuberculosis</em> isolates used for <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> benchmarking</strong></p> <p>The dataset comprises whole-genome sequence data published by <a href="https://doi.org/10.1016/S1473-3099(15)00062-6">Walker et al. 2015</a>. For the multiple sequence analysis, we proceeded as follows:</p> <ol> <li>Reads were downloaded from ENA BioProject <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA282721">PRJNA282721</a> (accessed on March 16<sup>th</sup>, 2023) and trimmed using Trimmomatic (<a href="https://pubmed.ncbi.nlm.nih.gov/24695404/">Bolger et al., 2014</a>) with <a href="https://github.com/B-UMMI/INNUca">INNUca</a> default settings;</li> <li>Quality-processed reads were individually mapped against the H37Rv reference genome (Genbank accession: <a href="https://www.ncbi.nlm.nih.gov/nuccore/NC_000962.3/">NC_000962.3</a>) using <a href="https://github.com/tseemann/snippy">Snippy</a> v4.5.1 and SNP-calling was performed on variant sites with the following criteria: a minimum proportion of reads differing from the reference of 70%, a minimum mapping quality of 30 and a minimum coverage for SNP calling of 10;</li> <li>A full alignment was extracted using Snippy’s core module (snippy-core), with masking of SNPs falling within known <em>M. tuberculosis</em> genomic regions with high GC content, repetitive elements and resistance-associated positions (corresponding to ~8% of the genome), as previously described for surveillance purposes (<a href="https://pubmed.ncbi.nlm.nih.gov/30948181/">Macedo et al., 2019</a>);</li> <li><em>M. tuberculosis </em>lineages were determined using tb-profiler v4.4.1 (<a href="https://pubmed.ncbi.nlm.nih.gov/31234910/">Phelan et al., 2019</a>), with samples from the <em>M. tuberculosis</em> complex other than <em>M. tuberculosis</em>, representing a mix of multiple lineages, or with less than 95% of mapped positions in the reference, being excluded;</li> <li>A filtered alignment comprising the maximum number of informative sites (88,562 nucleotide sites with at least one mutation in a given sequence) was extracted from the full alignment using the alignment_processing.py v1.1.0 (default settings) of <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a>, and then used as input for the benchmarking.</li> </ol> <p>In this repository, we provide two alignment files:</p> <ul> <li>Core_MTB_1787_strs.full.aln: this corresponds to the full multiple sequence alignment comprising 1787 samples and the reference (corresponding to the point 4 of the methodology).</li> <li>MTb_original_align_profile.fasta: this corresponds to the multiple sequence alignment comprising 1787 samples and the reference and only presenting the alignment informative sites (corresponding to the point 5 of the methodology)</li> </ul>
Translation Alignment: Ancient Greek to Latin. Annotation Style Guide and Gold Standard
<p>This dataset contains guidelines and a gold standard for the alignment of Ancient Greek texts with Latin scholarly translations. </p> <p>The gold standard consists of 100 fragments randomly selected from the <em>Digital Fragmenta Historicorum Graecorum </em>(DFHG) (https://www.dfhg-project.org/), which were aligned manually by Chiara Palladino and David J. Wright using Ugarit (https://ugarit.ialigner.com/). The Annotation Style Guide was developed for this project. The resulting Inter-Annotator-Agreement (IAA) is 90.5%. </p> <p>The materials available in this repository can be used to perform and evaluate alignments of various texts in Ancient Greek, to create gold standards, and to train automated translation alignment models. </p> <p>The Guidelines can be further adapted to address similar language pairs including inflected languages, or can provide a structure for the alignment of other historical texts against modern translations. However, the guidelines are not project-specific: they were specifically intended for the scenario of machine translation. Different research questions, such as translation history or pedagogy, may need further tweaking of these guidelines. </p> <p>For further information on Ugarit and translation alignment of historical languages, see http://ugarit.aligner.com/bib.php and follow us on Twitter (@ugarit_ty). <br> </p> <p> </p>
Virus alignments found in human cancer - aidinfo table
<p>The file contains one row for every analysis_id (column name 'aid') analyzed for the paper. An analysis_id is a unique TCGA BAM file.</p>
Do alignment and trimming methods matter for phylogenomic (UCE) analyses?
Alignment is a crucial issue in molecular phylogenetics because different alignment methods can potentially yield very different topologies for individual genes. But it is unclear if the choice of alignment methods remains important in phylogenomic analyses, which incorporate data from dozens, hundreds, or thousands of genes. For example, problematic biases in alignment might be multiplied across many loci, whereas alignment errors in individual genes might become irrelevant. The issue of alignment trimming (i.e. removing poorly aligned regions or missing data from individual genes) is also poorly explored. Here, we test the impact of 12 different combinations of alignment and trimming methods on phylogenomic analyses. We compare these methods using published phylogenomic data from ultraconserved elements (UCEs) from squamate reptiles (lizards and snakes), birds, and tetrapods. We compare the properties of alignments generated by different alignment and trimming methods (e.g., length, informative sites, missing data). We also test whether these datasets can recover well-established clades when analyzed with concatenated (RAxML) and species-tree methods (ASTRAL-III), using the full data (~5,000 loci) and subsampled datasets (10% and 1% of loci). We show that different alignment and trimming methods can significantly impact various aspects of phylogenomic datasets (e.g. length, informative sites). However, these different methods generally had little impact on the recovery and support values for well-established clades, even across very different numbers of loci. Nevertheless, our results suggest several "best practices" for alignment and trimming. Intriguingly, the choice of phylogenetic methods impacted the results most strongly, with concatenated analyses recovering significantly more well-established clades (with stronger support) than the species-tree analyses.
Data from: Properties of Markov chain Monte Carlo performance across many empirical alignments -- part I
<p>Nearly all current Bayesian phylogenetic applications rely on Markov chain Monte Carlo (MCMC) methods to approximate the posterior distribution for trees and other parameters of the model. These approximations are only reliable if Markov chains adequately converge and sample from the joint posterior distribution. While several studies of phylogenetic MCMC convergence exist, these have focused on simulated datasets or select empirical examples. Therefore, much that is considered common knowledge about MCMC in empirical systems derives from a relatively small family of analyses under ideal conditions. To address this, we present an overview of commonly applied phylogenetic MCMC diagnostics and an assessment of patterns of these diagnostics across more than 18,000 empirical analyses. Many analyses appeared to perform well and failures in convergence were most likely to be detected using the average standard deviation of split frequencies, a diagnostic that compares topologies among independent chains. Different diagnostics yielded different information about failed convergence, demonstrating that multiple diagnostics must be employed to reliably detect problems. The number of taxa and average branch lengths in analyses have clear impacts on MCMC performance, with more taxa and shorter branches leading to more difficult convergence. We show that the usage of models that include both Γ-distributed among-site rate variation and a proportion of invariable sites are not broadly problematic for MCMC convergence but are also unnecessary. Changes to heating and the usage of model-averaged substitution models can both offer improved convergence in some cases, but neither are a panacea.</p>
Data from: Properties of Markov chain Monte Carlo performance across many empirical alignments --part II
<p>Nearly all current Bayesian phylogenetic applications rely on Markov chain Monte Carlo (MCMC) methods to approximate the posterior distribution for trees and other parameters of the model. These approximations are only reliable if Markov chains adequately converge and sample from the joint posterior distribution. While several studies of phylogenetic MCMC convergence exist, these have focused on simulated datasets or select empirical examples. Therefore, much that is considered common knowledge about MCMC in empirical systems derives from a relatively small family of analyses under ideal conditions. To address this, we present an overview of commonly applied phylogenetic MCMC diagnostics and an assessment of patterns of these diagnostics across more than 18,000 empirical analyses. Many analyses appeared to perform well and failures in convergence were most likely to be detected using the average standard deviation of split frequencies, a diagnostic that compares topologies among independent chains. Different diagnostics yielded different information about failed convergence, demonstrating that multiple diagnostics must be employed to reliably detect problems. The number of taxa and average branch lengths in analyses have clear impacts on MCMC performance, with more taxa and shorter branches leading to more difficult convergence. We show that the usage of models that include both Γ-distributed among-site rate variation and a proportion of invariable sites are not broadly problematic for MCMC convergence but are also unnecessary. Changes to heating and the usage of model-averaged substitution models can both offer improved convergence in some cases, but neither are a panacea.</p>
Supplementary materials for "Relative Information Gain: Shannon entropy-based measure of the relative structural conservation in RNA alignments"
<p>Supplementary materials for "Relative Information Gain: Shannon entropy-based measure of the relative structural conservation in RNA alignments". These include precalculated RNA Blocks, MBRs (Matrix of Bear encoded RNA), sPSSMs (structural Position Specific Scoring Matrix), RIG (Relative Information Gain) scores, and plots calculated for 3016 Rfam 14.1 families. In particular:</p> <ul> <li><strong>alignments.zip:</strong> zipped file containing the structural alignments for each Rfam family.</li> <li><strong>RNA_Blocks.zip</strong>: zipped file containing the RNA blocks used to derive different substitution matrices.</li> <li><strong>MBRs.zip</strong>: zipped file containing the substitution matrices.</li> <li><strong>sPSSMs.zip</strong>: zipped file containing the structural Position Specific Scoring Matrices.</li> <li><strong>RIGs.zip</strong>: zipped file containing the RIG scores.</li> <li><strong>entropy.zip</strong>: zipped file containing the (rescaled) entropy.</li> <li><strong>plots.zip</strong>: zipped file containing the plots. </li> </ul> <p>All the scripts to build all these files are available at <a href="https://github.com/helmercitterich-lab/RIG">https://github.com/helmercitterich-lab/RIG</a>.</p>
Cilia density and flow velocity affect alignment of motile cilia from brain cells
<p>Here we store the supplementary Materials and Methods for the publication Cilia density and flow velocity affect alignment of motile cilia from brain cells.</p> <p>In the Supplementary methods we included additional information on the hydrodynamic simulations. </p> <p>Video1 and Video2 are videos referenced in the main text of the paper</p> <p>In the archive 'raw data and code.tar' , we provide raw images and codes to support the article.The complete dataset of raw images is more than 1 Tb. Here we are limited to 50Gb. The full dataset is available upon request.<br> <br> We choose to provide a full dataset of two culture at DIV 16, one treated with shear flow and a control without flow.</p> <p>For each of the two cultures, the videos with propelled particles are in the directory FL,<br> The bright field images without particles are stored in BF. Unfortunately we uploaded only few videos because of their large size. The results of the analysis of this dataset is reported in the directory analysis (available for each culture).</p> <p>Moreover we provide the code to analyse these data.<br> The analysis routine:</p> <p>Step 1: for each field of view (fov) getting the cilia beating direction from the FL images. This is done with PIV. The code is Step1_PIVanalysis.mat</p> <p>Step 2: for each fov getting ciliated cell position and CBF from the BF movies. Gather the cilia beating direction and cilia posion and frequency in a unique figure and matlab class (Res.mat). This is done in Step2_gatherResults.mat</p> <p>The results of these analysis are stored in the analysis folder for each culture.</p> <p>These routines are repeated for each experiment and results are then plotted to get trends. In the folder code4figures we report the code that we used to make the figures in the papers starting from a matlab file "all_results*.mat", where are gathered all the analysis.</p> <p>The code may improve in the future with more comments. please check Nicola's github page for the latest update. Please contact us for any problem. https://github.com/NicolaPellicciotta/Code4-Cilia-density-and-flow-velocity-affect-alignment-of-motile-cilia-from-brain-cells</p> <p>All the raw videos and code are in the archive.</p> <p> </p>
Multiple sequence alignments of sensor histidine kinases and response regulators
<p>The two FASTA files contain multiple sequence alignments of sensor histidine kinase and response regulator sequences. The source sequences were obtained by BLAST, clustered with usearch and aligned with muscle. More details to be found in Multamäki et al. 2021.</p>
Benchmark Database for Phonetic Alignments
<p>In the last two decades, alignment analyses have become an important technique in quantitative historical linguistics and dialectology. Phonetic alignment plays a crucial role in the identification of regular sound correspondences and deeper genealogical relations between and within languages and language families. Surprisingly, up to today, there are no easily accessible benchmark data sets for phonetic alignment analyses. Here we present a publicly available database of manually edited phonetic alignments which can serve as a platform for testing and improving the performance of automatic alignment algorithms. The database consists of a great variety of alignments drawn from a large number of different sources. The data is arranged in a such way that typical problems encountered in phonetic alignment analyses (metathesis, diversity of phonetic sequences) are represented and can be directly tested.</p>
Genome alignments for the article "Targeted reduction of highly abundant transcripts with pseudo-random primers"
<p>Sequence alignment (Moirai workflow management) for the article "Targeted reduction of highly abundant transcripts with pseudo-random primers". File names indicate unique run identifiers. In the manuscript, shorter names are used:</p> <ul> <li>NC12: NC12_1.CAGEscan_short-reads.20150629125015</li> <li>NC17: NC16-17_1.CAGEscan_short-reads.20150625154740</li> <li>NC22b: NC22b.CAGEscan_short-reads.20150625152335</li> <li>NCki: NCms10058_1.CAGEscan_short-reads.20150625154711</li> </ul>
On the alignment of academic publishers’ embargos with H2020 requirements - Dataset
<p> </p> <p>-- ATTENTION PLEASE: RIGHT NOW THE DATASET IS UNDER INDEPENDENT DOUBLE CHECK TO TEST THE PRESENCE OF POSSIBLE ERRORS : PLEASE CONTACT THE AUTHOR FOR FURTHER INFO --</p> <p> </p> <p>This dataset refers to the breifing paper "On the alignment of academic publishers’ embargos with H2020 requirements" (https://nexacenter.org/nexacenterfiles/WoS-Romeo-analysis%20APS-final.pdf) published in the ambit of the European project Pasteur4OA (http://www.pasteur4oa.eu/).</p> <p> </p> <p>We wanted to understand how the publishing behaviour of EU researchers might be affected by the H2020 policy requirement to ensure Open Access (OA) to all articles from EU-funded projects within 6 months for science and engineering projects and 12 months for humanities and social science studies. Many publishers impose an embargo on ‘Green’ Open Access, where researchers deposit their articles in repositories, and these embargoes can be longer than the maximum permitted by the H2020 policy. The issue was whether researchers may have to alter their publishing behaviour or can they continue to publish in journals of their choice and still comply with the H2020 requirements. The following research questions were posed:</p> <p>1) What is the level of compliance of current publishers’ embargo policies with the H2020 requirements?<br> 2) How many journals are compliant with the H2020 requirements?</p> <p>The overall findings were that 90% of publishers used by EU researchers to publish their work and 94% of articles published by EU researchers are compatible with the H2020 Open Access policy requirements. Our conclusion is that only in a small minority of cases – 5-10% – would EU researchers’ normal publishing behaviour run contrary to H2020 rules. </p> <p> </p> <p>The zip file containing the data used, divided by access to the post-prints: </p> <p>- ok Open Access</p> <p>- no Open Access</p> <p>- unclear</p> <p>- OA with restrictions, but compliant to H2020</p> <p>- OA with restrictions, not compliant to H200</p> <p>- OA with restrictions, not clear</p> <p>The data model is available on the paper (https://nexacenter.org/nexacenterfiles/WoS-Romeo-analysis%20APS-final.pdf)</p>
Raw data used for COI delineation of the Eupolybothrus species: Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar from: Eupolybothrus cavernicolus Komerički & Stoev sp. n. (Chilopoda: Lithobiomorpha: Lithobiidae): the first eukaryotic species description combining transcriptomic, DNA barcoding and micro-CT imaging data - Biodiversity Data Journal 1: e1013 (28 October 2013) https://doi.org/10.3897/BDJ.1.e1013
<p>Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar</p>
Integrated field-aligned radar data and analysis results
<p>Dataset used in "A statistical survey of heat input parameters into the cusp thermosphere" J. Geophys. Res. 2017, doi:10.1002/2016JA023594.</p>
Topical alignment in online social systems
<p>Data used in the work https://arxiv.org/abs/1707.06525.</p> <p>Published Files:</p> <p># Hashtag coocurrence graph <br> cooccurrence.txt —- each line: hashtag1 hashtag2 #cooccurrences</p> <p># Communities<br> communities.txt -- 1 line for each community</p> <p># Hashtags used by users<br> users_hashtags.json -- json file wherein, for each user, there is a list of hashtags and the number of times that he/she used it</p> <p># Edges <br> follow_friend.txt - edges from 2013</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.