Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,555
datasets available to search
ShareScore release 0.7.1
Dataset results
2,555 results for “catalogs”
FIGURE 15 in Illustrated and online catalog of type specimens of freshwater fishes in the Colección de Peces Dulceacuícolas of Instituto de Investigación de Recursos Biológicos Alexander von Humboldt (IAvH-P), Colombia
FIGURE 15. Paratype of Astroblepus nettoferreirai, IAvH-P 13209, 78.3 mm SL (scale bar = 1 cm). Photograph by C. DoNascimiento.
FIGURE 3 in Odontalgus dongbaiensis sp. n. from eastern China, and a world catalog of Odontalgini (Coleoptera: Staphylinidae: Pselaphinae)
FIGURE 3. Diagnostic features of Odontalgus dongbaiensis, male. A. Right antenna. B. Head dorsum, pronotum, and elytral base. C. Elytra and abdomen, in dorsal view. D. Protrochanter. E. Aedeagus, in dorsal view. F. Same, in lateral view. G. Same, in ventral view. Scale lines: A–C = 0.2 mm, D–G = 0.1 mm.
FIGURE 2 in Odontalgus dongbaiensis sp. n. from eastern China, and a world catalog of Odontalgini (Coleoptera: Staphylinidae: Pselaphinae)
FIGURE 2. Diagnostic features of Odontalgus dongbaiensis, male. A. Head, in dorsal view. B. Same, in lateral view, with maxillary palpomeres numbered. C. same, in ventral view. D. Pronotum. E. Prothorax, in lateral view. F. Prosternite. G. Left elytron. H. Meso- and metaventrite. Abbreviations: 1–4 indicate four discal striae, aldf—anterolateral discal foveae, amdfanteromedian discal foveae, bef—basal elytral foveae, ff—frontal fovea, gf—gular foveae, iblf—inner basolateral foveae, laflateral antebasal foveae, lmcf—lateral mesocoxal foveae, lmsf—lateral mesoventral foveae, lpcf—lateral procoxal foveae, lpp—lateral postantennal pits, maf—median antebasal fovea, mbf—mediobasal foveae, mmsf—median mesoventral foveae, mmtf—median metaventral foveae, oblf—outer basolateral foveae, pmmtf—posteromedian metaventral foveae, ss—sutural striae, vf—vertexal foveae. Scale lines: 0.2 mm.
FIGURES 1–9 in Catalog of the types of Coleoptera (Insecta) deposited at Museo de Historia Natural de la Universidad Nacional Mayor de San Marcos (MUSM), Lima, Peru
FIGURES 1–9. Type specimens of Coleoptera in the MUSM. 1, Dactylozodes (Nelsonozodes) huaylasiensis Moore & Diéguez (Buprestidae); 2, Trechisibus (Trechisibiodes) valenciai Etonti & Mateu (Carabidae); 3, Eumolpus viriditarsis viriditarsis Špringlová (Chrysomelidae); 4, Coleomegilla occulta González-Fuentes (Coccinellidae); 5, Lixus delatus Kuschel (Curculionidae); 6, Bystus decorator Leschen & Carlton (Endomychidae); 7, Epimetopus mendeli Fikáček, Barclay & Perkins (Hydrophilidae); 8, Lybanodes stigmatus Skelley in Skelley, Leschen & McHugh (Erotylidae); 9, Euspilotus (Hesperosaprinus) excavata Arriagada in Degallier, Arriagada, Kanaar, Moura, Tishechkin, Caterino & Warner (Histeridae). Holotypes (2, 6, 8); Paratypes (1, 3, 4, 5, 7, 9).
FIGURES 10–18 in Catalog of the types of Coleoptera (Insecta) deposited at Museo de Historia Natural de la Universidad Nacional Mayor de San Marcos (MUSM), Lima, Peru
FIGURES 10–18. Type specimens of Coleoptera in the MUSM. 10, Cratomorphus frankeae Bohórquez (Lampyridae); 11, Lagochile brunnea tenaensis Soula (Melolonthidae); 12, Pocadius maquipucunensis Leschen & Carlton (Nitidulidae); 13, Eurysternus inca Génier (Scarabaeidae); 14, Philothalpus lucieae Asenjo & Ribeiro-Costa (Staphylinidae); 15, Pseudopsis monica Asenjo & Ribeiro-Costa (Staphylinidae); 16, Hyperaspis esmeraldas Gordon & González-Fuentes (Coccinellidae); 17, Anthonomus onerosus Clark (Curculionidae); 18, Macrohaltica crypta Santisteban (Chrysomelidae). Holotypes (10, 14, 15, 17); Paratypes (11, 12, 13, 16, 18).
Enhanced seismicity catalog and output files of geomechanical modeling in the changning shale gas field
Open the record for dataset details and reuse information.
SERPENTINE catalog of solar cycle 25 multi-spacecraft solar energetic particle events
<p>This catalog contains multi-spacecraft solar energetic particle (SEP) events, which were observed with the new heliospheric spacecraft fleet in solar cycle 25. It has been created within the European Union’s Horizon 2020 project <em>SERPENTINE</em> (Solar energetic particle analysis platform for the inner heliosphere). The catalog comprises key SEP characteristics observed by five different observer locations as provided by <em>Solar Orbiter, Parker Solar Probe, STEREO A, Wind</em> and <em>SOHO</em> (at the Lagrangian point 1), and <em>BepiColombo</em>. It focuses on large events, which show energetic proton increases above 25 MeV observed at least at two spacecraft. The catalog provides not only key parameters of the proton event but also the same parameters for 1 MeV and 100 keV electrons, respectively.</p> <p>For more information, and if you use this catalog, please refer to the corresponding publication:</p> <blockquote> <p>The solar cycle 25 multi-spacecraft solar energetic particle event catalog of the SERPENTINE project<br>N. Dresing, A. Yli-Laurila, S. Valkila, J. Gieseler, D. E. Morosan, G. U. Farwa, Y. Kartavykh, C. Palmroos, I. Jebaraj, S. Jensen, P. Kühl, B. Heber, F. Espinosa, R. Gómez-Herrero, E. Kilpua, V.-V. Linho, P. Oleynik, L. A. Hayes, A. Warmuth, F. Schuller, H. Collier, H. Xiao, E. Asvestari, D. Trotta, J. G. Mitchell, C. M. S. Cohen, A. W. Labrador, M. E. Hill and R. Vainio<br>A&A, 687 (2024) A72<br>DOI: <a href="https://doi.org/10.1051/0004-6361/202449831">10.1051/0004-6361/202449831</a>, arXiv:<a href="https://arxiv.org/abs/2403.00658">2403.00658</a>. </p> </blockquote> <p>The csv files provided here are archived versions of the dynamic catalog available at <a href="https://data.serpentine-h2020.eu/catalogs/sep-sc25/">https://data.serpentine-h2020.eu/catalogs/sep-sc25/</a></p> <p>Provided are two semicolon-separated csv files containing the same data. The only difference is that the file <em>sep-sc25_extended_header.csv </em>includes an extensive header.</p> <ul> <li>Field descriptions <ul> <li>id: ID</li> <li>science_case: Science case</li> <li>date: Event date [UTC]</li> <li>flare_time: Flare time [UTC]</li> <li>flare_lat: Flare Carrington latitude [deg]</li> <li>flare_lon: Flare Carrington longitude [deg]</li> <li>flare_class: Flare class (GOES)</li> <li>flare_comments: Flare Comments</li> <li>radio_type2: Radio type II bursts</li> <li>decametric_type2_start: Decametric type II burst start time [UT]</li> <li>decametric_type2_stop: Decametric type II burst end time [UT]</li> <li>decametric_type2_freq_range: Decametric type II frequency range [MHz]</li> <li>radio_type2_start: Metric radio type II burst start time [UT]</li> <li>radio_type2_stop: Metric radio type II burst end time [UT]</li> <li>radio_type2_freq_range: Frequency range of metric type II burst [MHz]</li> <li>radio_imaging_available: Metric radio imaging available</li> <li>radio_comments: Radio comments</li> <li>cme_id: Associated CMEs</li> <li>solar_mach_link: Solar-Mach link</li> </ul> </li> </ul> <ul> <li>S/C codes <ul> <li>BepiC: BepiColombo</li> <li>L1: L1 (SOHO/Wind)</li> <li>PSP: Parker Solar Probe</li> <li>STA: STEREO A</li> <li>SolO: Solar Orbiter</li> </ul> </li> </ul> <ul> <li>S/C related field descriptions <ul> <li>{sc}_sc_lat: S/C Carrington latitude [deg]</li> <li>{sc}_sc_lon: S/C Carrington longitude [deg]</li> <li>{sc}_dist: S/C radial distance [au]</li> <li>{sc}_ip_shock_id: Associated IP shocks</li> <li>{sc}_p25MeV_onset_date: S/C protons 25 MeV onset date [UTC]</li> <li>{sc}_p25MeV_onset_time: S/C protons 25 MeV onset time [UTC]</li> <li>{sc}_p25MeV_onset_time_formatted: S/C protons 25 MeV onset time [UTC] (Formatted)</li> <li>{sc}_p25MeV_onset_averaging: S/C protons 25 MeV averaging used for onset [min]</li> <li>{sc}_p25MeV_onset_sector: S/C protons 25 MeV sector used for onset</li> <li>{sc}_p25MeV_peak_date: S/C protons 25 MeV peak date [UTC]</li> <li>{sc}_p25MeV_peak_time: S/C protons 25 MeV peak time [UTC]</li> <li>{sc}_p25MeV_peak_time_formatted: S/C protons 25 MeV peak time [UTC] (Formatted)</li> <li>{sc}_p25MeV_peak_flux: S/C protons 25 MeV peak flux [cm^-2 s^-1 sr^-1 MeV^-1]</li> <li>{sc}_p25MeV_peak_flux_formatted: S/C protons 25 MeV peak flux [cm^-2 s^-1 sr^-1 MeV^-1] (Formatted)</li> <li>{sc}_p25MeV_peak_averaging: S/C protons 25 MeV averaging used for peak [min]</li> <li>{sc}_p25MeV_peak_sector: S/C protons 25 MeV sector used for peak</li> <li>{sc}_p25MeV_injection_date: S/C protons 25 MeV inferred injection date [UTC]</li> <li>{sc}_p25MeV_injection_time: S/C protons 25 MeV inferred injection time [UTC]</li> <li>{sc}_p25MeV_pathlength: S/C protons 25 MeV path length used for inferred injection time [au]</li> <li>{sc}_p25MeV_sw_speed: S/C protons 25 MeV onset solar wind speed [km/s]</li> <li>{sc}_p25MeV_comments: S/C protons 25 MeV comments</li> <li>{sc}_e100keV_onset_date: S/C electrons 100 keV onset date [UTC]</li> <li>{sc}_e100keV_onset_time: S/C electrons 100 keV onset time [UTC]</li> <li>{sc}_e100keV_onset_time_formatted: S/C electrons 100 keV onset time [UTC] (Formatted)</li> <li>{sc}_e100keV_onset_averaging: S/C electrons 100 keV averaging used for onset [min]</li> <li>{sc}_e100keV_onset_sector: S/C electrons 100 keV sector used for onset</li> <li>{sc}_e100keV_peak_date: S/C electrons 100 keV peak date [UTC]</li> <li>{sc}_e100keV_peak_time: S/C electrons 100 keV peak time [UTC]</li> <li>{sc}_e100keV_peak_time_formatted: S/C electrons 100 keV peak time [UTC] (Formatted)</li> <li>{sc}_e100keV_peak_flux: S/C electrons 100 keV peak flux [cm^-2 s^-1 sr^-1 MeV^-1]</li> <li>{sc}_e100keV_peak_flux_formatted: S/C electrons 100 keV peak flux [cm^-2 s^-1 sr^-1 MeV^-1] (Formatted)</li> <li>{sc}_e100keV_peak_averaging: S/C electrons 100 keV averaging used for peak [min]</li> <li>{sc}_e100keV_peak_sector: S/C electrons 100 keV sector used for peak</li> <li>{sc}_e100keV_injection_date: S/C electrons 100 keV inferred injection date [UTC]</li> <li>{sc}_e100keV_injection_time: S/C electrons 100 keV inferred injection time [UTC]</li> <li>{sc}_e100keV_pathlength: S/C electrons 100 keV path length used for inferred injection time [au]</li> <li>{sc}_e100keV_sw_speed: S/C electrons 100 keV onset solar wind speed [km/s]</li> <li>{sc}_e100keV_comments: S/C electrons 100 keV comments</li> <li>{sc}_e1MeV_onset_date: S/C electrons 1 MeV onset date [UTC]</li> <li>{sc}_e1MeV_onset_time: S/C electrons 1 MeV onset time [UTC]</li> <li>{sc}_e1MeV_onset_time_formatted: S/C electrons 1 MeV onset time [UTC] (Formatted)</li> <li>{sc}_e1MeV_onset_averaging: S/C electrons 1 MeV averaging used for onset [min]</li> <li>{sc}_e1MeV_onset_sector: S/C electrons 1 MeV sector used for onset</li> <li>{sc}_e1MeV_peak_date: S/C electrons 1 MeV peak date [UTC]</li> <li>{sc}_e1MeV_peak_time: S/C electrons 1 MeV peak time [UTC]</li> <li>{sc}_e1MeV_peak_time_formatted: S/C electrons 1 MeV peak time [UTC] (Formatted)</li> <li>{sc}_e1MeV_peak_flux: S/C electrons 1 MeV peak flux [cm^-2 s^-1 sr^-1 MeV^-1]</li> <li>{sc}_e1MeV_peak_flux_formatted: S/C electrons 1 MeV peak flux [cm^-2 s^-1 sr^-1 MeV^-1] (Formatted)</li> <li>{sc}_e1MeV_peak_averaging: S/C electrons 1 MeV averaging used for peak [min]</li> <li>{sc}_e1MeV_peak_sector: S/C electrons 1 MeV sector used for peak</li> <li>{sc}_e1MeV_injection_date: S/C electrons 1 MeV inferred injection date [UTC]</li> <li>{sc}_e1MeV_injection_time: S/C electrons 1 MeV inferred injection time [UTC]</li> <li>{sc}_e1MeV_pathlength: S/C electrons 1 MeV path length used for inferred injection time [au]</li> <li>{sc}_e1MeV_sw_speed: S/C electrons 1 MeV onset solar wind speed [km/s]</li> <li>{sc}_e1MeV_comments: S/C electrons 1 MeV comments</li> <li>{sc}_ep_ratio: Ratio of Electrons (~1MeV) / Protons (25-40 MeV)</li> </ul> </li> </ul> <div> </div> <div><strong>CHANGELOG:</strong></div> <div> <ul> <li>2025-06-05 <ul> <li>Updated peak fluxes and peak times of PSP 1 MeV electrons, as well as PSP's e/p ratios (the previous flux values are erroneous!)</li> </ul> </li> <li>2024-09-06 <ul> <li>Flare information has been updated</li> <li>Updated cme_id info corresponding to updated CME catalog</li> <li>(Zenodo version only: change file format from "semicolon separated" to "comma separated, with fields containing commas surrounded by quotes")</li> </ul> </li> <li>2024-05-24<br> <ul> <li>Events 35, 36: flare times (GOES) were incorrectly provided as peak times. They have been changed to start times (like for all flares).</li> <li>All 8 events where SOLO/STIX has been used for flare determination had been time shifted to the Sun. Now they are shifted to 1 AU to be consistent with the GOES times.</li> </ul> </li> </ul> </div>
Numerical Sheet for "Station-orientation catalog for Australian broadband seismic stations"
<p>This XLSX file includes the numerical values for Tables S1-S3.</p> <p>This XLSX file includes the station orientation catalog for all stations estimated in Tarumi & Yoshizawa (2024, submitted to Seismica) (Sheet Tab1: Permanent stations excluding the S1 network; Sheet Tab2: the S1 network stations; Sheet Tab3: Temporary stations). The first column is an index number; the second includes the network code and station name. The third and fourth columns represent the period (start and end time) for calculating two statistical values (the median and mean) of station orientation (in the fifth and sixth columns, respectively) and its standard errors (seventh columns). The last column indicates the number of earthquakes used to estimate the misorientation during the designated period. All information is the same as in Tables S1–S3. </p> <p>"codesample.zip" includes the sample program (Jupyter notebook) and the sample data. </p>
A catalog of genes, genomes and species of the dog (Canis lupus familiaris) intestinal microbiota
<p></p><h1>Data sources</h1><br>This dataset was constructed using metagenomic sequencing data from the bioproject PRJEB20308 from Coelho et al. 2018 (129 samples)<br><h1>Metagenomic assembly</h1><br>First, sequencing adapters removal and read trimming was performed with fastp. Reads mapped on the host genome (ROS_Cfam_1.0 GCF_014441545.1) with bowtie2 were removed with samtools. Finally, Metagenomic assembly was performed with metaSPAdes. Contigs of less than 1500 bp were removed.<br><h1>MAGs recovery</h1><br>MAGs were generated with COMEBin (multi-coverage mode) and MAGs quality was assessed with CheckM2. MAGs with completeness < 70% or contamination > 5% or N50 < 5Kb were discarded. Pairwise Average Nucleotide Identity (ANI) was computed for all recovered MAGs with fastANI and dereplication at species level (ANI cutoff = 95%).<br><h1>Non-redundant gene catalog</h1><br>Genes were predicted on all contigs from metagenomic assemblies with Prodigal (parameters : -m -p meta). Genes were pooled and clustered with cd-hit-est (parameters -c 0.95 -aS 0.90 -G 0 -d 0 -M 0 -T 0) by choosing those from the longest contigs as representatives.<br><h1>MSPs recovery</h1><br>Reads were aligned against the non-redundant gene catalog with the Meteor software suite to produce a raw gene abundance table (1,0M genes quantified in 129 samples). Then, co-abundant genes were binned in 234 Metagenomic Species Pan-genomes (MSPs, i.e. gene clusters that likely belong to the same microbial species) using MSPminer.<br><h1>MAGs and MSPs taxonomic annotation</h1><br>Dereplicated MAGs were annotated with GTDB-Tk based on GTDB r220. Then, MAGs taxonomic annotation was propagated to the corresponding MSPs.<br><h1>Construction of the phylogenetic tree</h1><br>39 universal phylogenetic markers genes were extracted from the dereplicated MAGs with fetchMGs. Then, the markers were separately aligned with MUSCLE. The 40 alignments were merged and trimmed with trimAl (parameters: -automated1). Finally, the phylogenetic tree was computed with FastTreeMP (parameters: -gamma -pseudo -spr -mlacc 3 -slownni).<h1>Mapping rate distribution across public cohorts</h1>We generated mapping rate distribution plots using Meteor2 (default parameters), comparing performance between: PRJEB20308 (cohort used in catalogue assembly) and PRNJNA714112 (independent cohort not used in assembly).<p></p>
An updated catalog of genes and species of the pig gut microbiota
<p></p><h1>Dataset overview</h1><br>We built an updated catalog of 9.3M genes found in the pig gut microbiota.<br>Co-abundant genes were binned in 1523 Metagenomic Species Pan-genomes (MSPs) for which we provide taxonomic labels and a phylogenetic tree. In addition, we reconstituted 7059 Metagenome-Assembled Genomes (MAGs) covering of 760 Metagenomic Species and we extracted 7331 viral genomes from assemblies.<br>Finally, we used Pairwise Comparative Modelling to predict 6140 antibiotic resistance genes.<br><br>This dataset can be used to analyze shotgun sequencing data of the pig gut microbiota.<br><h1>Methods</h1><br><h2>Sequencing data availability</h2><br>Sequencing data from Xiao et al. (PRJEB11755, n=287) and Kim et al. (PRJEB32496, n=36) was downloaded from the European Nucleotide Archive.<br><h2>Sequencing data quality control</h2><br>Illumina adapters removal and read trimming was performed with fastp . Reads mapped on the host genome (GCF_000003025.6) with bowtie2 were removed with samtools.<br><h2>Metagenomic assembly</h2><br>Metagenomic assembly was performed with metaSPAdes. Contigs of less than 1500 bp were removed.<br><h2>MAGs creation</h2><br>Reads of each sample were aligned to their respective assembly with bowtie2 and results were indexed in sorted bam files with samtools. Then, contigs coverage was computed in each sample with jgi_summarize_bam_contig_depths. MAGs were generated with MetaBAT 2 and MaxBin2. Finally, results of both tools were combined with DAS Tool and MAGs quality was assessed with checkM. MAGs with completeness < 70% or contamination > 5% were discarded.<br><h2>Extraction of viral genomes</h2><br>Candidates viral sequences were identified in assemblies with VirFinder. Then, viral genomes quality was assessed with checkV and those low or undetermined quality were discarded.<br><h2>Non-redundant gene catalog</h2><br>Genes were predicted on all contigs with Prodigal (parameters : -m -p meta ). Genes with missing start codon or shorter than 99 bp were discarded.<br>Then, partial and complete genes were separately clustered with cd-hit-est (parameters -c 0.95 -aS 0.90 -G 0 -d 0 -M 0 -T 0 ). The two non-redundant gene sets were merged by considering at first complete genes from the longest contigs (contact us for futher details).<br><h2>MSPs creation</h2><br>Using the Meteor software suite, reads from each sample were mapped against the non redundant catalog to build a raw gene abundance table (9.3 million genes quantified in 323 samples). This table was submitted to MSPminer that reconstituted 1523 clusters of co-abundant genes named Metagenomic-Species Pangenomes (MSPs).<br>Quality control of each MSP was manually performed by visualizing heatmaps representative of the normalized gene abundance profiles.<br><h2>Taxonomic annotation</h2><br>MAGs and MSPs were annotated with GTDB-Tk based on GTDB Release 05-RS95.<br><h2>Construction of the phylogenetic tree</h2><br>39 universal phylogenetic markers genes were extracted from the 1523 MSPs (or the corresponding MAGs if available) with fetchMGs. Then, the markers were separately aligned with MUSCLE. The 40 alignments were merged and trimmed with trimAl (parameters: -automated1). Finally, the phylogenetic tree was computed with FastTreeMP (parameters: -gamma -pseudo -spr -mlacc 3 -slownni).<br><h2>Prediction of antibiotic resistance genes</h2><br>Antibiotic resistance genes were predicted with the Pairwise Comparative Modelling approach (last version available here).<h1>Mapping rate distribution across public cohorts</h1>We generated mapping rate distribution plots using Meteor2 (default parameters), comparing performance between: PRJEB11755 and PRJEB32496 (cohort used in catalogue assembly) and PRJEB60032 and PRJNA494875 (independent cohort not used in assembly).<p></p>
A catalog of genes and species of the human oral microbiota
<p></p><h1>Data sources</h1><br>The oral gene catalog was built using three primary sources:<br><br>Bacterial Genomes from the Human Oral Microbiome Database (HOMD).<br>Fungal Genomes from the NCBI RefSeq database.<br>Metagenomic Sequencing Data from multiple oral microbiome studies.<br>The creation of the oral gene catalog was a multi-step process, combining and refining genes from each source.<br><br><h1>Bacterial Genes</h1><br>A total of 1,505 bacterial genomes were downloaded from HOMD (version 20170215, accessed in December 2017). Genes shorter than 60 nucleotides or containing ambiguous bases were filtered out. Redundancy was removed using CD-HIT-EST (v4.6; parameters: -aS 0.9 -c 0.95 -T 0 -M 0 -t 0 -d 0 -G 0). This process yielded 1,459,394 unique HOMD genes for the catalog.<br><br><h1>Fungal Genes</h1><br>1,017 fungal genomes were downloaded from NCBI RefSeq (May 2017). For the 492 genomes lacking existing annotations, gene calling was performed using Genemark-ES in fungi mode. After initial redundancy removal with CD-HIT-EST (v4.6; parameters: -aS 0.9 -c 0.95 -T 0 -M 0 -t 0 -d 0 -G 0), genes were selected for inclusion only if their corresponding genome was present in at least 20% of the samples in one of the metagenomic cohorts, determined by mapping reads with Bowtie2 (v2.2.3). This led to the selection of 2,440,644 fungal genes.<br><br><h1>Metagenomic Sequencing Data</h1><br>The gene catalog was supplemented with data from 689 oral metagenomes, including newly sequenced samples, from the following studies:<br><br>Human Microbiome Project (HMP): 382 samples (bioproject PRJNA255439).<br>Chinese Cohort: 212 samples (bioproject PRJEB6997).<br>TwinsUK Cohort: 48 newly sequenced samples (bioproject: PRJEB38483).<br>Raw reads were subjected to quality control and trimmed using AlienTrimmer 0.4.0 (parameters: -k 10 -l 45 -m 5 -p 40 -q 20). Human sequences were removed by mapping against the human reference genome (GRCh38.p11) using Bowtie2 2.2.3. Metagenomic assembly was performed using SPAdes 3.9.0 (parameters: "-k 21,33,55 --only-assembler –meta" for Illumina paired-end data, or "--iontorrent -t 24 -m 300 -k 21,33,55 --only-assembler" for Ion Torrent single-end data). Contigs shorter than 500 bp or with coverage less than 2x were discarded. Gene calling was conducted with Prodigal (parameters: -m -p meta). Genes shorter than 60 bp were filtered out, and redundancy was removed with CD-HIT-EST (v4.6; parameters: -aS 0.9 -c 0.95 -T 0 -M 0 -t 0 -d 0 -G 0).<br><br><h1>Final Gene Catalog</h1><br>The final gene catalog was assembled by sequentially adding non-redundant genes from each data source. Genes from HOMD and fungal genomes were combined first using cd-hit-est-2d. Then, non-redundant genes from the HMP, Chinese, and TwinsUK cohorts were sequentially added using cd-hit-est-2d (same parameters as cd-hit-est). A final redundancy removal step was performed. This process resulted in a catalogue of 8.4 million non-redundant genes<br><br><h1>MSPs Recovery</h1><br>The 689 metagenomic samples were aligned against the final gene catalog using the Meteor software suite to produce a gene abundance table. Then, co-abundant genes were binned into 853 Metagenomic Species Pan-genomes (MSPs) using MSPminer.<br><br><h1>MSPs Taxonomic Annotation</h1><br>Taxonomic annotation for the MSPs was performed by aligning all core and accessory genes against representative genomes from the GTDB database (release r214) using blastn (task: megablast, word_size: 16).<br><br>A species-level assignment was given if over 50% of the genes matched a representative genome with a mean nucleotide identity of at least 95% and a mean gene length coverage of at least 90%.<br>The remaining MSPs were assigned to a higher taxonomic level (genus to superkingdom) if more than 50% of their genes shared the same annotation.<h1>Mapping rate distribution across public cohorts</h1>We generated mapping rate distribution plots using Meteor2 (default parameters), comparing performance between: PRJEB6997, PRJNA255439, PRJNA48479 (cohort used in catalogue assembly) and PRJEB24090, PRJEB28422, PRJEB45799 (independent cohort not used in assembly).<p></p>
A catalog of genes and species of the brown rat (Rattus norvegicus) gut microbiota
<p></p><h1>Dataset overview</h1><br>We built a catalog of 5.9M genes found in the brown rat gut microbiota. Co-abundant genes were binned in 1627 Metagenomic Species for which we provide taxonomic labels.<br><br>This dataset can be used to analyze shotgun sequencing data of the brown rat gut microbiota.<h1>Data sources </h1><br>Rat fecal (and milk) samples characterized by shotgun metagenomic sequencing during the Mamiprooffi project. Sequencing data will be submitted soon on the European Nucleotide Archive (Bioproject PRJEB57230)<br>The gene catalog of the Sprague-Dawley rat gut metagenome published by Pan et al.<br><h1>Metagenomic assembly</h1><br>Metagenomic assembly was performed on the Mamiprooffi samples (Data Source 1) with SPAdes (parameters: --iontorrent --careful). Contigs of less than 1500 bp or successfully aligned on the rat genome (Rnor_6.0) were removed.<br>Non-redundant gene catalog<br>Genes were predicted on all contigs with Prodigal (parameters : -m -p meta ). Genes with missing start codon or shorter than 99 bp were discarded.<br>Then, partial and complete genes were separately clustered with cd-hit-est (parameters -c 0.95 -aS 0.90 -G 0 -d 0 -M 0 -T 0 ). Finally, these two non-redundant gene sets were merged with the previously published catalog (Data Source 2) using cd-hit-est-2d by considering at first complete genes (contact us for futher details).<br>Functionnal annotation<br>KEGG Orthologs (KOs) were assigned to genes of the final catalog with KofamScan (version 1.3.0, KEGG 107 database)<br><h1>Metagenomic Species</h1><br>Using the Meteor software suite, reads from samples in Bioprojects PRJEB57230 and PRJEB22973 were mapped against the final non redundant catalog to build a raw gene abundance table (5.9 million genes quantified in 370 samples). This table was submitted to MSPminer and Canopy. A total of 1627 clusters of co-abundant genes or MetaGenomic Species (MGS) were discovered.<br>Quality control of each MGS was manually performed by visualizing heatmaps representative of the normalized gene abundance profiles.<br><h1>Taxonomic annotation of Metagenomic Species</h1><br>MGS taxonomic annotation was performed by aligning all core and accessory genes against the GTDB r214 representative genomes using blastn [4] (version 2.10.1, task = megablast, word_size = 16). The 20 best hits for each gene were kept. A species-level assignment was given if > 50% of the genes matched a GTDB representative genome with a mean identity ≥ 95% and mean gene length coverage ≥ 90%. The remaining MGS were assigned to a higher taxonomic levels (genus to superkingdom) if more than 50% of their genes had the same annotation.<h1>Mapping rate distribution across public cohorts</h1>We generated mapping rate distribution plots using Meteor2 (default parameters), comparing performance between: PRJEB22973 and PRJEB57230 (cohort used in catalogue assembly) and PRNJNA609596 (independent cohort not used in assembly).<p></p>
MIMIC2: Murine Intestinal Microbiota Integrated Catalog v2
<p></p><h1>Dataset overview</h1><br>The MIMIC2 dataset provides:<br>a non-redundant high-quality catalog of 5.0 million genes<br>6,967 Metagenome-Assembled Genomes (MAGs)<br>1,252 Metagenomic Species Pangenomes (MSPs)<br>This dataset can be used to analyze shotgun sequencing data of the murine gut microbiota.<br><h1>Methods</h1><br><h2>Data sources</h2><br>The MIMIC2 dataset was constructed using two different data sources:<br>Source 1: the Mouse Gastrointestinal Bacterial Catalogue (MGBC) which is a compilation of 276 genomes from cultured isolates and 45,218 metagenome-assembled genomes (MAGs) from 1,960 publicly available mouse metagenomes<br>Source 2: 68 samples of Messaoudene et al. (PRJNA783624) and 85 deeply sequenced samples from bioproject CNP0000619 published by Xiao et al.<br><h2>Metagenomic assembly</h2><br>De novo metagenomic assembly was performed on the 153 samples from the data Source 2. First, sequencing adapters removal and read trimming was performed with fastp. Reads mapped on the host genome (GCF_000001635.27) with bowtie2 were removed with samtools. Finally, Metagenomic assembly was performed with metaSPAdes. Contigs of less than 1500 bp were removed.<br><h2>MAGs recovery</h2><br>Reads of each sample from the data Source 2 were aligned to their respective assembly with bowtie2 and results were indexed in sorted bam files with samtools. Then, contigs coverage was computed in each sample with jgi_summarize_bam_contig_depths. MAGs were generated with MetaBAT 2 and MAGs quality was assessed with checkM. MAGs with completeness < 70% or contamination > 5% or N50 < 8Kb were discarded.<br><h2>Non-redundant gene catalog</h2><br>Genes were predicted on all contigs from the data Source 2 with Prodigal (parameters : -m -p meta ). Likewise, genes were predicted on all genomes from the data Source 1 (MGBC) with Prodigal (parameters : -m -p single ). Genes from the two data sources were pooled and those shorter than 90 bp or incomplete were discarded. Finally, genes were clustered with cd-hit-est (parameters -c 0.95 -aS 0.90 -G 0 -d 0 -M 0 -T 0 ) by choosing those from the longest contigs as representatives.<br><h2>MSPs recovery</h2><br>Samples from 19 cohorts (see below) were aligned against the non-redundant gene catalog with the Meteor software suite to produce a raw gene abundance table (5M genes quantified in 1374 samples). Then, co-abundant genes were binned in 1,252 Metagenomic Species Pan-genomes (MSPs, i.e. clusters of > 500 co-abundant genes that likely belong to the same microbial species) using MSPminer.<br><br>The 19 cohorts used to recover the MSPs are:<br>PRJNA783624<br>CNP0000619<br>PRJEB15095<br>PRJEB22007<br>PRJEB22710<br>PRJEB31298<br>PRJEB32790<br>PRJEB32890<br>PRJEB3374<br>PRJEB36943<br>PRJEB44286<br>PRJEB7759<br>PRJNA293255<br>PRJNA390686<br>PRJNA397886<br>PRJNA515074<br>PRJNA540893<br>PRJNA549182<br>PRJEB40719<br><h2>MSPs taxonomic annotation</h2><br>Representative genomes of the MMGC collection were annotated with GTDB-Tk based on GTDB r202. Then, taxonomic annotation of MMGC genomes was propagated to the corresponding MSPs.<br><br>For the MSPs without any corresponding MAG, taxonomic annotation was performed by alignment of all core and accessory genes against representative genomes of the GTDB database (release r202) using blastn (version 2.7.1, task = megablast, word_size = 16). A species-level assignment was given if > 50% of the genes matched the representative genome of a given species, with a mean nucleotide identity ≥ 95% and mean gene length coverage ≥ 90%. The remaining MSPs were assigned to a higher taxonomic level (genus to superkingdom), if more than 50% of their genes had the same annotation.<br><h2>Construction of the phylogenetic tree</h2><br>39 universal phylogenetic markers genes were extracted from the 1,252 MSPs (or the corresponding MAGs if available) with fetchMGs. Then, the markers were separately aligned with MUSCLE. The 40 alignments were merged and trimmed with trimAl (parameters: -automated1). Finally, the phylogenetic tree was computed with FastTreeMP (parameters: -gamma -pseudo -spr -mlacc 3 -slownni).<h1>Mapping rate distribution across public cohorts</h1>We generated mapping rate distribution plots using Meteor2 (default parameters), comparing performance between: CNP0000619, PRJEB15095, PRJEB22007, PRJNA783624 (cohort used in catalogue assembly) and PRJNA760892 (independent cohort not used in assembly).<p></p>
A catalog of genes and species of the human skin microbiota
<p></p><h1>Dataset overview</h1><br>This dataset provides:<br>a non-redundant high-quality catalog of 2.9 million genes<br>392 Metagenomic Species Pangenomes (MSPs)<br>This dataset can be used to analyze shotgun sequencing data of the human skin microbiota.<br><h1>How to use this dataset</h1><br>Create a gene abundance table by aligning reads from each sample against the catalog. For this purpose, you can use Meteor or NGLess. Then, normalize raw counts by gene length.<br>Taxonomic profiling: the abundance of each species can be estimated as the average abundance of its 100 first core genes. To reduce the false positive rate, only consider that a species is present if at least 10/100 marker genes are detected.<br><h1>Methods</h1><br><h2>Data sources</h2><br>This dataset was built using the following data sources:<br>118 isolate-derived genomes from the HMRGD<br>246 isolate-derived genomes from the Skin Microbial Genome Collection (SMGC)<br>1,407 skin metagenome assemblies from the Skin Microbial Genome Collection (SMGC)<br><h2>Non-redundant gene catalog</h2><br>After filtering out short contigs (<1500 bp), genes were predicted with Prodigal on genomes (mode: single) and metagenome assemblies (mode: meta). Complete genes (partial=00) were pooled and clustered with cd-hit-est (parameters -c 0.95 -aS 0.90 -G 0 -d 0 -M 0 -T 0) by choosing those from the longest contigs as representatives.<br><h2>Functional annotation</h2><br>KOs assignments were obtained with KofamScan using the KEGG 107 database.<br><h2>MSPs recovery</h2><br>Reads from the 1,120 skin metagenomes available in the bioproject PRJNA46333 were aligned against the non-redundant gene catalog with the Meteor software suite to produce a raw gene abundance table (2.9M genes quantified in 1,120samples). Then, co-abundant genes were binned in 392 Metagenomic Species Pan-genomes (MSPs, i.e. clusters of co-abundant genes that likely belong to the same microbial species) using MSPminer.<br><h2>MSPs taxonomic annotation</h2><br>Taxonomic annotation was performed by alignment of all core and accessory genes against representative genomes of the GTDB database (release r214) using blastn (version 2.7.1, task = megablast, word_size = 16). A species-level assignment was given if > 50% of the genes matched the representative genome of a given species, with a mean nucleotide identity ≥ 95% and mean gene length coverage ≥ 90%. The remaining MSPs were assigned to a higher taxonomic level (genus to superkingdom), if more than 50% of their genes had the same annotation.<br><h2>Construction of the phylogenetic tree</h2><br>39 universal phylogenetic markers genes were extracted from the MSPs (or the corresponding genome if available) with fetchMGs. Then, the markers were separately aligned with MUSCLE. The 40 alignments were merged and trimmed with trimAl (parameters: -automated1). Finally, the phylogenetic tree was computed with FastTreeMP (parameters: -gamma -pseudo -spr -mlacc 3 -slownni).<h1>Mapping rate distribution across public cohorts</h1>We generated mapping rate distribution plots using Meteor2 (default parameters), comparing performance between: PRJNA46333 (cohort used in catalogue assembly) and PRJEB80549 (independent cohort not used in assembly).<p></p>
FIGURES 82–85 in An annotated and illustrated catalog of the primary type material of Hemiptera deposited in the Florida State Collection of Arthropods
FIGURES 82–85. Dorsal habitus and labels of holotypes of 82). Gypona (Obtusana) ileota Freytag, 2005; 83). Gypona (Obtusana) toxum Freytag, 2005; 84). Gypona (Marganalana) woodruffi Freytag, 2005; 85). Ladoffa suttoni Freytag & Lozada, 2013. Scale bars = 1 mm.
FIGURES 43–45 in An annotated and illustrated catalog of the primary type material of Hemiptera deposited in the Florida State Collection of Arthropods
FIGURES 43–45. Dorsal habitus and labels of holotypes of 43). Parnisa santacruzensis Sanborn, 2019; 44). Pomponia daklakensis Sanborn, 2009; 45). Proarna gianucai Sanborn, 2008. Scale bars = 5 mm.
FIGURES 34–36 in An annotated and illustrated catalog of the primary type material of Hemiptera deposited in the Florida State Collection of Arthropods
FIGURES 34–36. Dorsal habitus and labels of holotypes of 34). Herrera nigratorquata Sanborn, 2018; 35). Herrera nigropercula Sanborn, 2020; 36). Herrera polygramma Sanborn, 2020. Scale bars = 5 mm.
FIGURES 40–42 in An annotated and illustrated catalog of the primary type material of Hemiptera deposited in the Florida State Collection of Arthropods
FIGURES 40–42. Dorsal habitus and labels of holotypes of 40). Herrera signifera Sanborn, 2019; 41). Neocicada centramericana Sanborn, 2005. Dorsal habitus and labels of neotype of 42). Neotibicen superbus (Fitch, 1855). Scale bars = 5 mm.
FIGURES 7–9 in An annotated and illustrated catalog of the primary type material of Hemiptera deposited in the Florida State Collection of Arthropods
FIGURES 7–9. Dorsal habitus and labels of holotypes of 7). Carineta acommosis Sanborn, 2020; 8). Carineta apicoinfuscata Sanborn, 2011; 9). Carineta digitata Sanborn, 2020. Scale bars = 5 mm.
FIGURES 52–54 in An annotated and illustrated catalog of the primary type material of Hemiptera deposited in the Florida State Collection of Arthropods
FIGURES 52–54. Dorsal habitus and labels of holotypes of 52). Procollina nuevoleonensis Sanborn, 2018; 53). Procollina stigmosa Sanborn, 2018; 54). Selymbria boliviaensis Sanborn, 2019. Scale bars = 5 mm.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.