Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
10
datasets available to search
ShareScore release 0.7.1
Dataset results
10 results for “MLST”
Supplemental material of the Streptococcus pyogenes whole genome MLST schema deposited in Chewie-NS
<p>This supplemental material includes the lists of accession numbers for the Blackwell et al. and NCBI RefSeq assemblies used to populate the whole genome MLST schema for <em>Streptococcus pyogenes</em>, the UniProt identifiers of the reference proteomes used for schema annotation and the set of complete genomes, and associated metadata, used for schema creation.</p> <p>The wgMLST schema was created with <a href="https://github.com/B-UMMI/chewBBACA">chewBBACA</a> and is publicly available at <a href="https://chewbbaca.online/species/1/schemas/1">chewie-NS</a>, where a more detailed description of schema creation, annotation and curation can be found.</p>
INNUENDO whole genome and core genome MLST schemas and datasets for Campylobacter jejuni
<p><strong>Dataset</strong></p> <p>Raw reads deposited in the European Nucleotide Archive (ENA) or in the NCBI Sequence Read Archive (SRA) as <em>C. jejuni</em> were retrieved in April 2017. In total 5,691 genomes passed the INNUca v3.1 pipeline have been selected. Additionally, 566 raw reads previously published in <a href="https://www.ncbi.nlm.nih.gov/pubmed/27041390">Kovanen et al., 2016</a>, <a href="https://www.ncbi.nlm.nih.gov/pubmed/28348829">Llarena et al., 2016</a>, <a href="https://www.ncbi.nlm.nih.gov/pubmed/25232158">Kovanen et al., 2014</a>, <a href="https://www.ncbi.nlm.nih.gov/pubmed/24655229">Kovanen et al., 2014</a> and <a href="http://www.sciencedirect.com/science/article/pii/S0740002016310449?via%3Dihub">Gacia-Sanchez et a., 2017</a> were included. The database also includes 269 <em>C. jejuni</em> belonging to the INNUENDO Sequence Dataset (<a href="https://www.ebi.ac.uk/ena/data/view/PRJEB27020">PRJEB27020</a>). Genomes were assembled using <a href="https://github.com/INNUENDOCON/INNUca">INNUca v3.1 pipeline</a> and passed the QC. </p> <p>File 'Metadata/Cjejuni_metadata.txt' contains metadata information for each strain including country and year of isolation, source classification and taxa of the host, classical pubMLST 7 genes ST and CC classification. </p> <p>The directory 'Genomes' contains all the 6,526 INNUca V3.1 assemblies of the strains listed in 'Metadata/Cjejuni_metadata.txt'.</p> <p><strong>Schema creation and validation</strong></p> <p>Draft genome assemblies were annotated using Prokka and initial pangenome was defined using Roary. The <a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA CreateSchema.py</em></a> was used for creating a whole genome schema starting from roary pangenome. The schema was initially composed by 5,447 loci and has been populated with the 6,526 <em>C. jejuni</em> genomes. The quality of the loci has been assessed using <a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA Schema Evaluation</em></a>. Loci with single alleles and those with high length variability (i.e. if more than 1 allele is outside the mode +/- 0.05 size) have been removed. The wgMLST schema has been further curated, excluding all those loci detected as “Repeated Loci” and loci annotated as “non-informative paralogous hit (NIPH/ NIPHEM)” or “Allele Larger/ Smaller than length mode (ALM/ ASM)” by the <a href="https://github.com/B-UMMI/chewBBACA/wiki/2.-Allele-Calling"><em>chewBBACA Allele Calling</em></a> engine in more than 1% of the <em>C. jejuni</em> genomes dataset.</p> <p>File 'Schema/Cjejuni_wgMLST_2795_schema.tar.gz' contains the wgMLST schema formatted for chewBBACA and includes a total of 2,795 loci.</p> <p>File 'Schema/Cjejuni_cgMLST_678_listGenes.txt' contains the list of genes from the wgMLST schema which defines the cgMLST schema. The cgMLST schema consists of 678 loci and has been defined as the loci present in at least the 99.9% of the 6,526 <em>C. jejuni</em> genomes. Genomes have no more than 2% of missing loci.</p> <p>File 'Allele_Profles/Cjejuni_wgMLST_alleleProfiles.tsv' contains the wgMLST allelic profile of the 6,526 <em>C. jejuni</em> genomes of the dataset. Please note that missing loci follow the annotation of chewBBACA Allele Calling software.</p> <p>File 'Allele_Profles/Cjejuni_cgMLST_alleleProfiles.tsv' contains the cgMLST allelic profile of the 6,526 <em>C. jejuni</em> genomes of the dataset. Please note that missing loci are indicated with a zero.</p> <p><strong>Additional citations</strong></p> <p>The schema are prepared to be used with <a href="https://github.com/B-UMMI/chewBBACA/wiki"><strong>chewBBACA</strong></a>. When using the schema in this repository please cite also</p> <blockquote> <p>Silva M, Machado M, Silva D, Rossi M, Moran-Gilad J, Santos S, Ramirez M, Carriço J. chewBBACA: A complete suite for gene-by-gene schema creation and strain identification. 15/03/2018. M Gen 4(3): doi:10.1099/mgen.0.000166 <a href="http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166">http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166</a></p> </blockquote>
INNUENDO whole genome and core genome MLST schemas and datasets for Escherichia coli
<p><strong>Dataset</strong></p> <p>As reference dataset, 2,218 public draft or complete genome assemblies and available metadata of <em>Escherichia coli</em> have been downloaded from <a href="https://enterobase.warwick.ac.uk/species/index/ecoli">EnteroBase</a> in April 2017. Genomes have been selected on the basis of the ribosomal ST (rST) classification available in <a href="https://enterobase.warwick.ac.uk/species/index/ecoli">EnteroBase</a>: from the same rST, genomes have been randomly selected and downloaded. The number of samples for each rST in the final dataset is proportional to those available in <a href="https://enterobase.warwick.ac.uk/species/index/ecoli">EnteroBase</a> in April 2017. The dataset includes also 119<em> </em>Shiga toxin-producing <em>E.coli</em> genomes assembled with<a href="https://github.com/B-UMMI/INNUca"> INNUca v3.1 </a>belonging to the INNUENDO Sequence Dataset (<a href="https://www.ebi.ac.uk/ena/data/view/PRJEB27020">PRJEB27020</a>).</p> <p>File 'Metadata/Ecoli_metadata.txt' contains metadata information for each strain including source classification, taxa of the hosts, country and year of isolation, serotype, pathotype, classical pubMLST 7 genes ST classification, assembly source/method and Enterobase barcode. </p> <p>The directory 'Genomes' contains the 119 INNUca v3.1 assemblies of the strains listed in 'Metadata/Ecoli_metadata.txt'. Enterobase assemblies can be downloaded from http://enterobase.warwick.ac.uk/species/ecoli/search_strains using 'barcode'.</p> <p><strong>Schema creation and validation</strong></p> <p>The wgMLST schema from <a href="https://enterobase.warwick.ac.uk/species/ecoli/download_data">EnteroBase</a> have been downloaded and curated using <a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA AutoAlleleCDSCuration</em></a> for removing all alleles that are not coding sequences (CDS). The quality of the remain loci have been assessed using <a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA Schema Evaluation</em></a> and loci with single alleles, those with high length variability (i.e. if more than 1 allele is outside the mode +/- 0.05 size) and those present in less than 0.5% of the <em>Escherichia</em> genomes in <a href="https://enterobase.warwick.ac.uk/species/index/ecoli">EnteroBase</a> at the date of the analysis (April 2017) have been removed. The wgMLST schema have been further curated, excluding all those loci detected as “Repeated Loci” and loci annotated as “non-informative paralogous hit (NIPH/ NIPHEM)” or “Allele Larger/ Smaller than length mode (ALM/ ASM)” by the <a href="https://github.com/B-UMMI/chewBBACA/wiki/2.-Allele-Calling"><em>chewBBACA Allele Calling</em></a> engine in more than 1% of a dataset composed by 2,337 <em>Escherichia coli</em> genomes.</p> <p>File 'Schema/Ecoli_wgMLST_7601_schema.tar.gz' contains the wgMLST schema formatted for chewBBACA and includes a total of 7,601 loci.</p> <p>File 'Schema/Ecoli_cgMLST_2360_listGenes.txt' contains the list of genes from the wgMLST schema which defines the cgMLST schema. The cgMLST schema consists of 2,360 loci and has been defined as the loci present in at least the 99% of the 2,337 <em>Escherichia coli</em> genomes. Genomes have no more than 2% of missing loci.</p> <p>File 'Allele_Profles/Ecoli_wgMLST_alleleProfiles.tsv' contains the wgMLST allelic profile of the 2,337 <em>Escherichia coli</em> genomes of the dataset. Please note that missing loci follow the annotation of chewBBACA Allele Calling software.</p> <p>File 'Allele_Profles/Ecoli_cgMLST_alleleProfiles.tsv' contains the cgMLST allelic profile of the 2,337 <em>Escherichia coli</em> genomes of the dataset. Please note that missing loci are indicated with a zero.</p> <p><strong>Additional citations</strong></p> <p>The schema are prepared to be used with <a href="https://github.com/B-UMMI/chewBBACA/wiki"><strong>chewBBACA</strong></a>. When using the schema in this repository please cite also:</p> <blockquote> <p>Silva M, Machado M, Silva D, Rossi M, Moran-Gilad J, Santos S, Ramirez M, Carriço J. chewBBACA: A complete suite for gene-by-gene schema creation and strain identification. 15/03/2018. M Gen 4(3): doi:10.1099/mgen.0.000166 <a href="http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166">http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166</a></p> </blockquote> <p><em>Escherichia coli</em> schema is a derivation of EnteroBase <em>E. coli</em> <a href="http://enterobase.warwick.ac.uk/">EnteroBase</a> wgMLST schema. When using the schema in this repository please cite also:</p> <blockquote> <p>Alikhan N-F, Zhou Z, Sergeant MJ, Achtman M (2018) A genomic overview of the population structure of <em>Salmonella</em>. PLoS Genet 14 (4):e1007261. <a href="https://doi.org/10.1371/journal.pgen.1007261">https://doi.org/10.1371/journal.pgen.1007261</a></p> </blockquote>
INNUENDO whole genome and core genome MLST schemas and datasets for Salmonella enterica
<p><strong>Dataset</strong></p> <p>As reference dataset, 4,307 public available draft or complete genome assemblies and available metadata of <em>Salmonella enterica</em> have been downloaded from public repositories (i.e. <a href="https://enterobase.warwick.ac.uk/">EnteroBase</a>, <a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information NCBI</a>and <a href="https://www.ebi.ac.uk/">The European Bioinformatics Institute EMBL-EBI</a>; accessed April 2017). The collection includes 1,465 <em>S.</em> Enteritidis, 2,442 <em>S.</em>Typhimurium, and 400 of other frequently isolated serovars in Europe. The dataset includes also 153 <em>S.</em>Typhimurium variant 4,[5],12:i:- collected from different Italian regions between 2012 and 2014 during a surveillance study and 129 <em>S.</em> Enteritidis belonging to the INNUENDO sequence dataset (<a href="https://www.ebi.ac.uk/ena/data/view/PRJEB27020">PRJEB27020</a>). The 282 additional genomes were assembled using <a href="https://github.com/B-UMMI/INNUca">INNUca v3.1</a>.</p> <p>File 'Metadata/Senterica_metadata.txt' contains metadata information for each strain including source classification, host taxa, year and country of isolation, serotype, classical pubMLST 7 genes ST classification, and source/method of the assembly. </p> <p>The directory 'Genomes' contains all the 4,589 assemblies of the strains listed in 'Metadata/Senterica_metadata.txt'. Please note that genomes marked as 'Enterobase' have been downloaded from Enterobase webpage http://enterobase.warwick.ac.uk.</p> <p><strong>Schema creation and validation</strong></p> <p>The wgMLST schema from <a href="https://enterobase.warwick.ac.uk/species/senterica/download_data">EnteroBase</a> have been downloaded and curated using <a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA AutoAlleleCDSCuration</em></a> for removing all alleles that are not coding sequences (CDS). The quality of the remain loci have been assessed using <a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA Schema Evaluation</em></a> and loci with single alleles, those with high length variability (i.e. if more than 1 allele is outside the mode +/- 0.05 size) and those present in less than 0.5% of the <em>Salmonella</em> genomes in <a href="https://enterobase.warwick.ac.uk/species/index/senterica">EnteroBase</a> at the date of the analysis (April 2017) have been removed. The wgMLST schema have been further curated, excluding all those loci detected as “Repeated Loci” and loci annotated as “non-informative paralogous hit (NIPH/ NIPHEM)” or “Allele Larger/ Smaller than length mode (ALM/ ASM)” by the <a href="https://github.com/B-UMMI/chewBBACA/wiki/2.-Allele-Calling"><em>chewBBACA Allele Calling</em></a> engine in more than 1% of a dataset composed by 4,589 <em>Salmonella</em> genomes.</p> <p>File 'Schemas/Senterica_wgMLST_ 8558_schema.tar.gz' contains the wgMLST schema formatted for chewBBACA and includes a total of 8,558 loci.</p> <p>File 'Schemas/Senterica_cgMLST_ 3255_listGenes.txt' contains the list of genes from the wgMLST schema which defines the cgMLST schema. The cgMLST schema consists of 3,255 loci and has been defined as the loci present in at least the 99% of the 4,589 <em>Salmonella</em> genomes. Genomes have no more than 2% of missing loci.</p> <p>File 'Allele_Profles/Senterica_wgMLST_alleleProfiles.tsv' contains the wgMLST allelic profile of the 4,589 <em>Salmonella</em> genomes of the dataset. Please note that missing loci follow the annotation of chewBBACA Allele Calling software.</p> <p>File 'Allele_Profles/Senterica_cgMLST_alleleProfiles.tsv' contains the cgMLST allelic profile of the 4,589 <em>Salmonella</em> genomes of the dataset. Please note that missing loci are indicated with a zero.</p> <p><strong>Additional citations</strong></p> <p>The schema are prepared to be used with <a href="https://github.com/B-UMMI/chewBBACA/wiki"><strong>chewBBACA</strong></a>. When using the schema in this repository please cite also:</p> <blockquote> <p>Silva M, Machado M, Silva D, Rossi M, Moran-Gilad J, Santos S, Ramirez M, Carriço J. chewBBACA: A complete suite for gene-by-gene schema creation and strain identification. 15/03/2018. M Gen 4(3): doi:10.1099/mgen.0.000166 <a href="http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166">http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166</a></p> </blockquote> <p><em>Salmonella enterica</em> schema is a derivation of EnteroBase <em>Salmonella </em><a href="http://enterobase.warwick.ac.uk/">EnteroBase</a> wgMLST schema. When using the schema in this repository please cite also:</p> <blockquote> <p>Alikhan N-F, Zhou Z, Sergeant MJ, Achtman M (2018) A genomic overview of the population structure of <em>Salmonella</em>. PLoS Genet 14 (4):e1007261. <a href="https://doi.org/10.1371/journal.pgen.1007261">https://doi.org/10.1371/journal.pgen.1007261</a></p> </blockquote>
INNUENDO whole genome and core genome MLST schemas and datasets for Yersinia enterocolitica
<p><strong>Dataset</strong></p> <p>All the raw reads deposited in the European Nucleotide Archive (ENA) or in the NCBI Sequence Read Archive (SRA) as <em>Y. enterocolitica</em> at the time of the analysis (August 2018) were retrieved using <a href="https://github.com/B-UMMI/getSeqENA">getSeqENA</a>. A total of 252 genomes were successfully assembled using <a href="https://github.com/B-UMMI/INNUca">INNUca v3.1</a>. In addition to public available genomes, the database includes 79 novel <em>Y. enterocolitica</em> strains which belong to the INNUENDO Sequence Dataset (<a href="https://www.ebi.ac.uk/ena/data/view/PRJEB27020">PRJEB27020</a>). </p> <p>File 'Metadata/Yenterocolitica_metadata.txt' contains metadata information for each strain including country and year of isolation, source classification, taxon of the host, serotype, biotype, pathotype (according to patho_typing software) and classical pubMLST 7 genes ST according to <a href="http://jcm.asm.org/content/53/1/35.long">Hall et al., 2005</a>. </p> <p>The directory 'Genomes' contains all the 331 INNUca V3.1 assemblies of the strains listed in 'Metadata/Yenterocolitica_metadata.txt'.</p> <p><strong>Schema creation and validation</strong></p> <p>All the 331 genomes were used for creating the schema using <strong><a href="https://github.com/B-UMMI/chewBBACA">chewBBACA suite</a></strong>. The quality of the loci have been assessed using <a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA Schema Evaluation</em></a> and loci with single alleles, those with high length variability (i.e. if more than 1 allele is outside the mode +/- 0.05 size) and those present in less than 1% of the genomes have been removed. The wgMLST schema have been further curated, excluding all those loci detected as “Repeated Loci” and loci annotated as “non-informative paralogous hit (NIPH/ NIPHEM)” or “Allele Larger/ Smaller than length mode (ALM/ ASM)” by the <a href="https://github.com/B-UMMI/chewBBACA/wiki/2.-Allele-Calling"><em>chewBBACA Allele Calling</em></a> in more than 1% of a dataset.</p> <p>File 'Schema/Yenterocolitica_wgMLST_ 6344_schema.tar.gz' contains the wgMLST schema formatted for chewBBACA and includes a total of 6,344 loci.</p> <p>File 'Schema/Yenterocolitica_cgMLST_ 2406_listGenes.txt' contains the list of genes from the wgMLST schema which defines the cgMLST schema. The cgMLST schema consists of 2,406 loci and has been defined as the loci present in at least the 99% of the 331<em> Y. enterocolitica</em> genomes. Genomes have no more than 2% of missing loci.</p> <p>File 'Allele_Profles/Yenterocolitica_wgMLST_alleleProfiles.tsv' contains the wgMLST allelic profile of the 331 <em>Y. enterocolitica</em> genomes of the dataset. Please note that missing loci follow the annotation of chewBBACA Allele Calling software.</p> <p>File 'Allele_Profles/Yenterocolitica_cgMLST_alleleProfiles.tsv' contains the cgMLST allelic profile of the 331 <em>Y. enterocolitica</em> genomes of the dataset. Please note that missing loci are indicated with a zero.</p> <p><strong>Additional citation</strong></p> <p>The schema are prepared to be used with <a href="https://github.com/B-UMMI/chewBBACA/wiki"><strong>chewBBACA</strong></a>. When using the schema in this repository please cite also:</p> <blockquote> <p>Silva M, Machado M, Silva D, Rossi M, Moran-Gilad J, Santos S, Ramirez M, Carriço J. chewBBACA: A complete suite for gene-by-gene schema creation and strain identification. 15/03/2018. M Gen 4(3): doi:10.1099/mgen.0.000166 <a href="http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166">http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166</a></p> </blockquote>
A core genome MLST (cgMLST) scheme for Staphylococcus pseudintermedius
<p>A core-genome MLST scheme for <em>Staphylococcus pseudintermedius </em>was developed using chewBBACA version 2.8.5 (Silva et al, 2018) with default settings (https://github.com/B-UMMI/chewBBACA_tutorial). A training file was generated using Prodigal version 2.6.3 (Hyatt et al, 2010). The scheme was generated using 74 complete genomes and is comprised of 1,356 genes present in 99% of the genomes. cgMLST types were designated relative to conventional MLST sequence types. Phylogenetic trees from the chewBBACA allele calls were constructed using GrapeTree version 1.5.0 and the RapidNJ algorithm (Zhou et al, 2018).</p> <p>Due to format updates, two versions of the scheme are included: one usable with chewBBACA version 2.8.5, one updated for the most recent version in September 2024, version 3.3.10.</p> <p><strong>References:</strong></p> <p>- Hyatt D, Chen GL, LoCascio PF, Land ML, Larimer FW, Hauser LJ. Prodigal: Prokaryotic Gene Recognition and Translation Initiation Site Identification. BMC Bioinformatics 2010;11:119. https://doi.org/10.1186/1471-2105-11-119. </p> <p><span>- Silva M, Machado MP, Silva DN, Rossi M, Moran-Gilad J, Santos S, et al. chewBBACA: A complete suite for gene-by-gene schema creation and strain identification. <em>Microb Genom</em> 2018;4:</span><span> </span>e000166<span> 10.1099/mgen.0.000166.</span></p> <p>- Zhou Z, Alikhan NF, Sergeant MJ, Luhmann N, Vaz C, Francisco AP, et al. GrapeTree: visualization of core genomic relationships among 100,000 bacterial pathogens. <em>Genome Res</em> 2018;28(9):1395-404 10.1101/gr.232397.117.</p> <p> </p>
Data from: Extending RAD tag analysis to microbial ecology: a comparison between multi locus sequence typing (MLST) and 2b-RAD to investigate Listeria monocytogenes genetic structure
The advent of next-generation sequencing (NGS) has dramatically changed bacterial typing technologies, increasing our ability to differentiate bacterial isolates. Despite it is now possible to sequence a bacterial genome in a few days and at reasonable costs, most genetic analyses do not require whole-genome sequencing, which also remains impractical for large population samples due to the cost of individual library preparation and bioinformatics. More traditional sequencing approaches, however, such as MultiLocus Sequence Typing (mlst) are quite laborious and time-consuming, especially for large-scale analyses. In this study, a genotyping approach based on restriction site-associated (RAD) tag sequencing, 2b-RAD, was applied to characterize Listeria monocytogenes strains. To verify the feasibility of the method, an in silico analysis was performed on 30 available complete genomes. For the same set of strains, in silico mlst analysis was conducted as well. Subsequently, 2b-RAD and mlst analyses were experimentally carried out on 58 isolates collected from food samples or food-processing sites. The obtained results demonstrate that 2b-RAD predicts mlst types and often provides more detailed information on population structure than mlst. Moreover, the majority of variants differentiating identical sequence type isolates mapped against accessory fragments, thus providing additional information to characterize strains. Although mlst still represents a reliable typing method, large-scale studies on molecular epidemiology and public health, as well as bacterial phylogenetics, population genetics and biosafety could benefit of a low cost and fast turnaround time approach such as the 2b-RAD analysis proposed here.
Video demonstration of the Standing to table behaviour element in the Minimally interactive group (behaviour element ID code: MLST) for Korcsok and Korondi (2023), Biologia Futura
<p>The video demonstrates the behaviour element: Standing to table in the minimally interactive experimental group, as exhibited by a social robot (behaviour element ID code: <strong>MLST</strong>). The video is part of an ethogram cataloguing the behaviour elements of the robot, described in the publication: <em><strong>How do you do the things that you do? - Ethological approach to the description of robot behaviour</strong></em> submitted to Biologia Futura (2023) by Korcsok, B. and Korondi, P.</p>
Data from: Extending RAD tag analysis to microbial ecology: a comparison between multi locus sequence typing (MLST) and 2b-RAD to investigate Listeria monocytogenes genetic structure
Open the record for dataset details and reuse information.
Expanded MLST genotyping and comparative genomic hybridization of Campylobacter coli isolates from multiple hosts
GEO Series GSE16787. Campylobacter coli. 132 samples. Type: Genome variation profiling by array.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.