Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,666
datasets available to search
ShareScore release 0.9.0
Dataset results
1,666 results for “human genome”
Genomic DNA transposition induced by human PGBD5 (Accompanying Scripts for Computational Analyses)
<p>We include the bash scripts, python code and command-line parameters used to prepare, map, and analyze NGS sequencing data generated from a modified version of flanking sequence exponential anchored (FLEA) PCR to identify genomic insertion locations of a transposable element reporter construct. In summary, we map these reads to a hybrid genome consisting in the human reference hg19 and the reporter plasmid, identify reads that span the insertion breakpoint and thus recover the genomic insertion loci. </p> <p>Please find more details in the supplied README. </p> <p> </p> <p>This collection of scripts corresponds to the data analysis in the following publication:</p> <p>Genomic DNA transposition induced by human PGBD5</p> <p>Anton Henssen, Elizabeth Henaff, Eileen Jiang, Amy Eisenberg, Julianne R. Carson, Camila Villasante, Mondira Ray, Eric Still, Melissa Burns, Jorge Gandara, Cedric Feschotte, Christopher E. Mason, Alex Kentsis</p> <p> </p> <p> </p>
Tandem repeat catalog of the human genome generated from long-read assemblies
<p>Allele sequences of polymorphic loci (VCF) and README for all version 2 (2.0 + 2.1) files</p>
Genomic footprints of (pre) colonialism: Population declines in urban and forest túngara frogs coincident with historical human activity
<p>Urbanisation is rapidly altering ecosystems, leading to profound biodiversity loss. To mitigate these effects, we need a better understanding of how urbanisation impacts dispersal and reproduction. Two contrasting population demographic models have been proposed that predict that urbanisation either promotes (facilitation model) or constrains (fragmentation model) gene flow and genetic diversity. Which of these models prevails likely depends on the strength of selection on specific phenotypic traits that influence dispersal, survival, or reproduction. Here, we a priori examined the genomic impact of urbanisation on the Neotropical túngara frog (<em>Engystomops pustulosu</em>s), a species known to adapt its reproductive traits to urban selective pressures. Using whole-genome resequencing for multiple urban and forest populations we examined genomic diversity, population connectivity and demographic history. Contrary to both the fragmentation and facilitation models, urban populations did not exhibit substantial changes in genomic diversity or differentiation compared to forest populations, and genomic variation was best explained by geographic distance rather than environmental factors. Adopting an a posteriori approach, we additionally found both urban and forest populations to have undergone population declines. The timing of these declines appears to coincide with extensive human activity around the Panama Canal during the last few centuries rather than recent urbanisation. Our study highlights the long-lasting legacy of past anthropogenic disturbances in the genome and the importance of considering the historical context in urban evolution studies as anthropogenic effects may be extensive and impact non-urban areas on both recent and older timescales. </p>
The structure of simple satellite variation in the human genome and its correlation with centromere ancestry (Supplemental Data)
<p>Accompanying <a href="https://github.com/is-the-biologist/1KGP_SATS" target="_blank" rel="noopener">Github</a></p> <p><strong>Supplemental File 1.</strong> BLAST results of k-mer concatemers against T2T-CHM13-v2.0.</p> <p><strong>Supplemental File 2.</strong> Annotations of centromeres, and telomeres of T2T-CHM13-v2.20. Table of abundance of k-mers in annotated regions as numpy file from BLAST hits. Abundance of k-mers across genome in 100kb bins from BLAST hits as .npz files accessible by example:</p> <p> import numpy as np<br> dense = np.load("filename.npz")<br> dense["chr1"]<br> <br><strong>Supplemental File 3</strong>. Table of pairwise R2 between simple satellites and table of pairwise interspersion OR between simple satellites. Folder containing QQ plots of negative binomial fit of satellite copy number distribution used to qualitatively asses model fit.</p> <p><strong>Supplemental File 4. </strong>Materials and results of cenGRM analysis. Boundaries used for centromeric regions of each cenGRM, cenGRMs in GCTA format, and tables with the results of cenGRM GCTA runs. Also provide pdfs of the dendrograms/heatmaps produced from UPGMA clustering of each cenGRM. </p> <p><strong>Supplemental File 5</strong> Non-human significant BLAST hits from BLAST-ing k-mer concatamers to non-human sequences.</p> <p><strong>Supplemental Table 1.</strong> Copy number normalized to 1x depth given GC bias of 126 most abundant satellites analyzed in paper in each individual. Additional columns represent metadata of the individual:</p> <ul> <li>instrument: sequencer instrument name used to sequence library.</li> <li>run: sequencer run of the library.</li> <li>flow: flowcell ID of the ibrary.</li> <li>pop: 1,000 Genomes Project population ID.</li> <li>superpop: 1,000 Genomes Project superpopulation ID.</li> <li>reads: average autosomal read depth of the library.</li> </ul> <p><strong>Supplemental Table 2. </strong>Copy number normalized to 1x depth given GC bias of the top 126 most abundant satellites analyzed in paper in each individual of the 1KGP, plus estimates of the same satellites in CHM13 short-read libraries subsampled from 18x-0.5x, 18x depth simulated library of the T2T-CHM13v2.0 assembly analyzed using k-Seek, and Tandem Repeat Finder results of the T2T-CHM13v2.0 asembly <a href="https://doi.org/10.1126/science.abk3112" target="_blank" rel="noopener">Hoyt 2022</a>.</p> <p><strong>Supplemental Table 3.</strong> Copy number normalized to 1x depth given GC bias of all tandem repeats with k-mer <= 20 (6,309) found collectively in the CHM13 short-read libraries subsampled from 18x-0.5x, 18x depth simulated library of the T2T-CHM13v2.0 assembly analyzed using k-Seek, and Tandem Repeat Finder results of the T2T-CHM13v2.0 asembly <a href="https://doi.org/10.1126/science.abk3112" target="_blank" rel="noopener">Hoyt 2022</a>.</p>
ATAC-seq processing resources for the GRCh38 (hg38) assembly of the human genome
<p>A collection of publicly available, but preprocessed, reference data for the analysis of ATAC-seq samples using the GRCh38 (hg38) assembly of the human genome using the <a href="https://doi.org/10.5281/zenodo.6323634">Ultimate ATAC-seq Data Processing & Analysis Pipeline</a> (details in the documentation on GitHub).</p>
Genome-wide expression profile of AAV2-infected normal human fibroblasts
<p>Datasets containing the genome-wide expression profile of AAV2-infected normal human fibroblasts.</p> <p>Raw data: results--A1--over--M1-3.txt</p> <p>p<0.01, reads>40: A1_M1_p0.01_r40_fc not restricted.txt</p>
Comparative genomics of human distal lung Streptococci
<p><strong>Comparative genomics of human lung streptococcal isolates</strong></p> <p>1. All analysis pipelines and scripts are on the GitHub page of Slipa Kanungo: <a href="https://github.com/slipa17/Whole-genome-sequencing-and-comparative-genomics-of-human-lung-streptococcal-isolates">https://github.com/slipa17/Whole-genome-sequencing-and-comparative-genomics-of-human-lung-streptococcal-isolates</a></p> <p>2. Additional analysis pipelines (especially Dataset S12) are on the GitHub page of Garance Sarton-Lohéac: <a href="https://github.com/gsartonl/Publication_Sarton-Loheac_2022">https://github.com/gsartonl/Publication_Sarton-Loheac_2022</a></p> <p>3. All raw data were uploaded to NCBI SRA BioProject <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1001255">PRJNA1001255</a></p> <p><strong>Supplementary Table bundle: for peer review purposes</strong></p> <p><strong>Supplementary Datasets</strong></p> <ul> <li><strong>Dataset S1_Lung_Streptococcus_genomes_metaQUAST: </strong>MetaQUAST (Quality Assessment Tool for Metagenome Assemblies) output including HTML and PDF reports, summary statistics including total contigs, assembly size, and N50. Coverage analysis assesses how well reference genomes are represented, contig length distribution plots visualize contig length ranges, mis-assembly analysis detects potential errors and graphical representations to visualize assemblies.</li> <li><strong>Dataset S2_Lung_streptococcus_isolate_genomes: </strong>Nucleotide FASTA files of six lung streptococcal isolates obtained that were obtained via whole genome sequencing.</li> <li><strong>Dataset S3_Lung_isolates_genome_annotation_prokka: </strong>Output folders after annotation of six lung streptococcal isolates with PROKKA. This includes protein FASTA, GenBank files and GFF annotations.</li> <li> <p><strong>Dataset S4_TYGS_dDDH_analysis: </strong>Contains results of TYGS analysis from DSMZ including downloadable reports. Outputs including taxonomic identification with genus, species, and strain details, a TYGS index for tracking genomes, genome quality assessment metrics, GBDP whole genome and 16S rRNA phylogenetic tree files, comparisons with reference type strains in the TYGS database with table.</p> </li> <li> <p><strong>Dataset S5_Reference_type_strains_TYGS_genomes: </strong>Nucleotide FASTA files of 47 closely related reference <em>Streptococcus</em> genomes listed by TYGS and downloaded from NCBI.</p> </li> <li> <p><strong>Dataset S6_Reference_type_strains_TYGS_proteins: </strong>Protein FASTA files of 47 closely related reference <em>Streptococcus</em> genomes listed by TYGS and downloaded from NCBI. </p> </li> <li> <p><strong>Dataset S7_Lung_streptococcus_isolate_proteins: </strong>Protein FASTA files of 6 six lung streptococcal isolates. </p> </li> <li> <p><strong>Dataset S8_OrthoFinder_core_genome: </strong>OrthoFinder is a bioinformatics tool that offers comprehensive outputs for orthology inference across multiple genomes. The output includes overall statistics, gene duplication information, orthologous genes, orthologous gene tree, single copy orthologous genes and STAG evolutionary trees.</p> </li> <li> <p><strong>Dataset S9_Pan-Strep_BLAST_db: </strong>BLAST database using the <code>makeblastdb</code> command of NCBI datasets command line tool. This is constructed using protein FASTA files of 47 closely related reference <em>Streptococcus</em> genomes listed by TYGS and downloaded from NCBI.</p> </li> <li><strong>Dataset S10_OrthoVenn_cluster_files: </strong>OrthoVenn is a web-based tool for orthologous gene comparison. Downloadable results include Venn diagrams depicting shared and unique orthologous clusters amongst species, tabular results detailing genes within each cluster and their annotations. Functional enrichment analysis for Gene Ontology terms and KEGG pathways are also provided. enhances biological insights.</li> <li> <p><strong>Dataset S11_COG_analysis: </strong>Results of COG analysis of six lung streptococcal isolates individually using eggNOG (evolutionary genealogy of genes: Non-supervised Orthologous Groups) webtool. The output includes information on Clusters of Orthologous Groups (COGs) categorizing them into functional groups such as metabolism, information storage and processing, and cellular processes and signalling.</p> </li> <li> <p><strong>Dataset S12_CAZymes_lung_streptococci: </strong>Results of CAZyme analysis using a custom rule-based pipeline mostly based on dbCAN (Database for Carbohydrate-Active enZymes) provides information on the carbohydrate-active enzymes present in genomic datasets. The output includes the annotation of enzymes involved in the degradation, modification, or biosynthesis of carbohydrates: glycoside hydrolases (GH), glycosyltransferases (GT), carbohydrate-binding modules (CBM), Auxillary Activities (AA), Carbohydrates Esterases (CE) and Polysaccharide lyases (PL).</p> </li> <li> <p><strong>Dataset S13_pneumolysin_analysis: </strong>Results alignment and phylogeny of Pnuemolysin protein in <em>Streptococcus pneumoniae</em>, <em>Streptococcus pseudopneumoniae</em> and Streptococcus isolate P2E5 found by ABRIcate analysis. Visual plots by pyGenomeViz.</p> </li> <li> <p><strong>Dataset S14_capsule_analysis:</strong> Results from BLAST analysis of <em>Streptococcus pneumoniae </em>D39 capsular biosynthesis operon genes against the Pan-Strep (Dataset S10). Extracted of matching genes followed alignment and phylogeny. Visual plots by pyGenomeViz.</p> </li> <li> <p><strong>Dataset S15_Lung_isolate_HOMD_TYGS_comparison: </strong>Protein FASTA files of 47 closely related reference <em>Streptococcus</em> genomes listed by TYGS and downloaded from NCBI, 6 six lung streptococcal isolates and 47 streptococcal genomes from downloaded from human oral microbiome database (eHOMD).</p> </li> </ul>
Data for paper MethPhaser: methylation-based haplotype phasing of human genomes
<p>Data for paper MethPhaser: methylation-based haplotype phasing of human genomes. </p> <p>Files start with R9 or R10 are HG002 sample data.</p> <p>Zip files are block connection intermediate for MethPhaser. </p> <p>GTFs are block assignments. </p>
A comprehensive catalog of approxiamte short tandem repeat regions on autosomes and sex chromosomes of the human genome GRCh38
<p>To obtain a general TR catalog across the human genome, we identified genomic intervals with a stretch of approximate repetitions of a DNA motif ranging from 1-6bp on GRCh38 autosomes and sex chromosomes by using STRfinder (v1.0), and each STR region was annotated based on gencode.V38 (https://www.gencodegenes.org/human/release_38.html). To end up, we successfully found 1,656,159 TR intervals, covering 1.107653% (34.2 Mbp) of GRCh38 (https://console.cloud.google.com/storage/browser/_details/genomics-public-data/resources/broad/hg38/v0/Homo_sapiens_assembly38.fasta). </p>
Human intestinal Bacteria Collection (HiBC): Genome sequences
<p>The <a href="https://hibc.rwth-aachen.de/" target="_blank" rel="noopener">Human intestinal Bacteria Collection (HiBC)</a> is a collection of bacterial strains, isolated from the human gut for which 16S rRNA gene sequences, genome sequences and culture conditions are made available to the research community. In addition to previously described bacteria, we include strains that represent novel species which have been taxonomically described and validly named, or will be in the future. This collection will be updated regularly.</p> <p>This dataset includes the genome sequences of the isolates in the FASTA nucleotide format. Plasmids sequences when present are located at the very end of the file.</p> <p><strong>UPDATE v3</strong>: The genome of one of our isolate had been unfortunately swapped. This mistake has been now corrected on Zenodo and Coscine. The genome of <em>Segatella sinensis</em> CLA-AA-H117 should be considered correct with 103 contigs and 3 671 232 nt. Please note that the genome available at the NCBI is the correct one (GCA_040324585.2). Two typos regarding taxonomy have been corrected as well: <em>Maccoya intestinihominis</em> has been corrected to <em>Maccoyia intestinihominis</em> and <em>Faecousia faecis</em> to <em>Faecousia intestinalis</em>.</p>
hg19KIndel: Ethnicity normalized human reference genome
<p>The above zip files (hg19KIndel_Resource.zip) contains the following folders. Specific Details about how each file within the below mentioned folders were derived are present in individual README files for each folder</p> <p>1) hg19Kindel</p> <p> -hg19Kindel.fa - fasta file representing the modified assembly</p> <p>2) Gene Annotations</p> <p> -hg19_refGene.txt - RefSeq gene annotation downloaded from UCSC table browser (for hg19)<br> -hg19Kindel_refGene.txt - RefSeq gene annotation corresponding to hg19Kindel Genome</p> <p>3) SnpEff_Database</p> <p> -snpEffectPredictor.bin - Binary file used by SnpEff to annotate variants</p> <p>4) LiftOver_and_Chain_File</p> <p> -hg19Kindeltohg19.over.chain - UCSC chain file for lifting coordinates from hg19Kindel to hg19<br> -convert_cordinates_vcf.py - Python script to liftover variants(vcf) (only point coordinates) called on hg19Kindel to hg19 coordinate frame</p> <p> </p>
A human gut Faecalibacterium prausnitzii fatty acid amide hydrolase- Genome Annotation
<p>Genome annotation for <em>F. prausnitzii </em>Bg7063</p> <p><em>Science </em><strong>386</strong>, eado6828 (2024)</p> <p>DOI: 10.1126/science.ado6828</p> <p>Undernutrition in Bangladeshi children is associated with disruption of postnatal gut microbiota assembly; compared with standard therapy, a microbiota-directed complementary food (MDCF) substantially improved their ponderal and linear growth. Here, we characterize a fatty acid amide hydrolase (FAAH) from a growth-associated intestinal strain of Faecalibacterium prausnitzii cultured from these children. This enzyme, expressed and purified from Escherichia coli, hydrolyzes a variety of N-acylamides, including oleoylethanolamide (OEA), neurotransmitters, and quorum sensing N-acyl homoserine lactones; it also synthesizes a range of N-acylamides, notably N-acyl amino acids. Treating germ-free mice with N-oleoylarginine and N-oleolyhistidine, major products of FAAH OEA metabolism, markedly affected expression of intestinal immune function pathways. Administering MDCF to Bangladeshi children considerably reduced fecal OEA, a satiety factor whose levels were negatively correlated with abundance and expression of their F. prausnitzii FAAH. This enzyme, structurally and catalytically distinct from mammalian FAAH, is positioned to regulate levels of a variety of bioactive molecules.</p> <p> </p>
A unified genealogy of modern and ancient genomes: Unified, inferred tree sequences of 1000 Genomes, Human Genome Diversity, and Simons Genome Diversity Projects with ancient samples
<p>Unified, inferred tree sequences built from the 1000 Genomes phase 3, Human Genome Diversity, and Simons Genome Diversity Projects with high coverage sequenced ancient samples. The ancient samples are the Altai, Chagyrskaya, and Vindija Neanderthals, the Denisovan, and a high-coverage family of four from the Afanasievo Culture.</p> <p>Each tree sequence is the arm of an autosome (the short arm of acrocentric chromosomes are not included). Tree sequences were inferred with <a href="https://tsinfer.readthedocs.io/">tsinfer</a> version 0.2.1 and <a href="https://tsdate.readthedocs.io/en/latest/">tsdate</a> version 0.1.4, as described in <a href="http://www.biorxiv.org/content/10.1101/2021.02.16.431497v2">Wohns et al. (2021)</a>. The files were compressed using <a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. All data is in GRCh38.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on <a href="https://github.com/awohns/unified_genealogy_paper">GitHub</a>. A description can be found in the Supplementary Material of <a href="https://www.biorxiv.org/content/10.1101/2021.02.16.431497v2">Wohns et al. (2021)</a>.</p> <p>Tree sequences can be decompressed as follows:</p> <pre><code>$ tsunzip hgdp_tgp_sgdp_high_cov_ancients_chr1_p.dated.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed in Python using <a href="https://tskit.readthedocs.io/">tskit</a>. </p> <pre><code>import tskit ts = tskit.load("hgdp_tgp_sgdp_high_cov_ancients_chr1_p.dated.trees") # ts is an instance of tskit.TreeSequence print("The short arm of chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Accessing variant sites in the tree sequence provides the position and id of variants:</p> <pre><code>import json site = ts.site(1000) site_metadata = json.loads(site.metadata) print("The position of site 1000 is {} and its ID is {}.".format(site.position, site_metadata["ID"]))</code></pre> <p>Metadata associated with individuals and populations was derived from the original sources (<a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/working/20130606_sample_info/20130606_g1k.ped">TGP</a>, <a>HGDP</a>, and <a href="https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/SGDP_metadata.279public.21signedLetter.samples.txt">SGDP</a>) and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code>ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code>pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p>
Metagenome assembled genome database of a human cohort and fecal reactors
<p><strong>HumanCohort_annotations.tsv.zip:</strong> This is the custom MAG database (n=2447 MAGs) and corresponding annotations that were used in Borton 2022: "Targeted curation of the gut microbial gene content modulating human cardiovascular disease". The citation will be updated upon publication of the manuscript. Metagenome assembled genomes were generated from fecal metagenomes derived from a 54 person cohort and anoxic methylated amine enrichments. </p> <p><strong>HumanCohortmetabolism_summary.xlsx.zip: </strong> This is the annotation summary for 2447 MAGs in the cohort database. </p> <p><strong>Quality_Abundance_CohortMAGs.xlsx: </strong>This is a genome inventory of the 2447 MAGs in the cohort database including genome statistics and relative abundance. </p> <p><strong>orig_1D_NMR_fids.zip: </strong>NMR data derived from anoxic methylated amine enrichments. </p>
Human respiratory genome catalogue
<p>Human respiratory microbiome plays a crucial role in respiratory health, but there lack a comprehensive respiratory genome catalogue to study the microbiome. In this study, we collected whole-metagenome shotgun sequencing data of 4,067 samples, and sequenced long reads of 124 samples, leading 9.08 Tbp short-read and 0.42 Tbp long-read data, respectively. The data, integrated with a novel assembly algorithm, enables us to obtain a comprehensive human respiratory genome catalogue (RGC). This RGC displays high quality, with 190,443 contigs over one Kbps and an N50 length exceeding 13 Kbps; it comprises 159 high-quality and 393 medium-quality genomes, including 117 previously uncharacterized respiratory bacteria. To facilitate respiratory microbiome study, we present three important archives for RGC applications, including genome sequences, gene sequences in GBK format, taxonomical profile, ARG profile, VG profile, and MGC profile.</p>
A human genome editing-based MLL-AF4 acute lymphoblastic leukemia model recapitulates key cellular and molecular leukemogenic features. (Processed data)
<p>The prognosis of infant B-cell acute lymphoblastic leukemia (iB-ALL) remains dismal, especially in patients harboring the MLL-AF4 (KTM2A-AFF1) rearrangement, which arises prenatally in early hematopoietic stem/progenitor cells (HSPCs) and accounts for 80% of iB-ALL and 10% of non-infant cases. MLL-AF4+ B-ALL shows a bimodal localization of the MLL gene breakpoint within the MLL break cluster region, and two subgroups of patients based on the gene expression pattern of the HOXA/MEIS cluster have been identified. The pathogenic mechanisms in MLL- AF4+ B-ALL are challenging to study functionally due to the absence of faithful human cellular models recapitulating the disease phenotype and latency. Here, we assess the molecular contribution and leukemogenic capacity of MLL breakpoints occurring in either intron 10 (MLL i10 , centromeric) or intron 12 (MLL i12 , telomeric) in ontogenically-different human HSPCs sourced prenatally (fetal liver) and neonatally (cord blood). CRISPR-Cas9-induced MLL-AF4 (MA) targeting either MLL i10 (M i10 A) or MLL i12 (M i12 A) causes MA-driven in vitro myeloid immortalization in both fetal liver- and cord blood-CD34+ HSPCs. The centromeric location of the MLL breakpoint, but not the cellular ontogeny, determined the expression of HOXA/MEIS1 genes in MLL-edited cells. Centromeric MLL breakpoints endowed enhanced myeloid clonogenic replating to MLL- edited CD34+ HSPCs. The cellular ontogeny and the location of the MLL breakpoint also influenced the capacity of MLL-edited CD34+ HSPCs to initiate pro-B-ALL in vivo, which faithfully recapitulated the molecular, transcriptomic and methylome profiles of patients with primary MA+ iB-ALL. Our data provide key insights into the cellular and molecular leukemogenic determinants of MA+ iB-ALL. This dataset contains processed RNAseq and DNA methylation data from the abovementioned study.</p>
Data from: Were bed bugs the first urban pest insect? Genome-wide patterns of bed bug demography mirror global human expansion
Open the record for dataset details and reuse information.
Genomic footprints of (pre) colonialism: Population declines in urban and forest túngara frogs coincident with historical human activity
Open the record for dataset details and reuse information.
Detection and analysis of complex structural variation in human genomes across populations and in brains of donors with psychiatric disorders
Open the record for dataset details and reuse information.
Negative linkage disequilibrium between amino acid changing variants reveals interference among deleterious mutations in the human genome
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.