Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
20,966
datasets available to search
ShareScore release 0.9.0
Dataset results
20,966 results for “family”
Gene family data from the PhyloGenes (release version 5.0, phylogenes.org)
<p>The data files were generated from the PhyloGenes 5.0release (see release notes <a href="https://conf.arabidopsis.org/display/PHGSUP/About+PhyloGenes">here</a>).</p> <p>About the two zip files: </p> <p>1. phyloXML.zip</p> <p>PhyloGenes gene family trees in PhyloXML format, one file per family (e.g. <family_ID>.xml).</p> <p>The following information is provided for each node of a tree:<br>1) leaf node:<br>branch length<br>name <gene_id><br>taxonomy scientific_name<br>sequence accession <UniProt ID></p> <p>2) non-leaf node:<br>branch length<br>events <duplication or speciation></p> <p><br>2. panther_csv.zip</p> <p>Functional information of family members in CSV format, one file per family (e.g. <family_ID>.csv). </p> <p>A CSV file includes the following columns:<br>Uniprot ID<br>Gene <Gene name. If none then Gene ID><br>Gene ID<br>Gene name<br>Organism<br>Subfamily name</p> <p>The columns displayed after 'Subfamily name', if any, are GO annotations. Each column is a GO molecular function or biological process term that is annotated to at least one member of the gene family AND the annotation is supported by an experimental evidence (indicated by 'EXP') or phylogenetic inference (indicated by 'IBA'). A '0' indicates absence of either annotations.</p>
Phlorest phylogeny derived from Kolipakam et al. 2018 'A Bayesian phylogenetic study of the Dravidian language family'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Kolipakam V, Jordan FM, Dunn M, Greenhill SJ, Bouckaert R, Gray RD & Verkerk A. 2018. A Bayesian phylogenetic study of the Dravidian language family. R. Soc. Open Sci. 5: 171504.</p> </blockquote>
Phlorest phylogeny derived from Bouckaert et al. 2012 'Mapping the Origins and Expansion of the Indo-European Language Family'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Bouckaert RR, Lemey P, Dunn M, Greenhill SJ, Alekseyenko AV, Drummond AJ, Gray RD, Suchard MA & Atkinson QD. 2012. Mapping the Origins and Expansion of the Indo-European Language Family. Science, 337(6097), 957-960.</p> </blockquote>
Phlorest phylogeny derived from Birchall et al. 2016 'A combined comparative and phylogenetic analysis of the Chapacuran language family'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Birchall, Joshua, Michael Dunn, and Simon J. Greenhill. 2016. A combined comparative and phylogenetic analysis of the Chapacuran language family. International Journal of American Linguistics 82 (3): 255–84. doi: 10.1086/687383</p> </blockquote>
Alpha-Galactosaminidase family GH191 protein from Environmental sample (99.2% identity to Myxococcus fulvus enzyme): X-ray diffraction images
<p><span>This submission includes a zip archive of diffraction images recorded with the Dectris EIGER X 9M detector at the DIAMOND beamline I04-1. The model of the crystal structure and associated information can be found in the Protein Data Bank entry 9EP5. This is a case of crystal pathology – partial disorder. The model has C 2 2 21 symmetry and two molecules per asymmetric unit with occupancies 1 and 1/3. The molecule with partial occupancy overlaps with a symmetry related molecule.</span></p>
Proprotein convertase subtilisin/kexin type 9 (PCSK9) inhibitor therapy reduces the level of DNA damage in patients with heterozygous familial hypercholesterolemia
<p><span><span>Heterozygous Familial Hypercholesterolaemia (HeFH) is a common autosomal dominant genetic disease (1:300) characterized by elevated LDL-C leading to premature atherosclerosis. Treatment with a PCSK9 inhibitor (iPCSK9) is recommended in high cardiovascular risk FH patients if the treatment goal is not achieved on maximal tolerated statin plus ezetimibe. </span></span><span>The aim of this study was to </span><span><span>examine the changes in DNA damage in HeFH patients associated with iPCSK9 use. </span></span><span>Fifty-seven patients were included: a normolipidemic group (control; n=20) and patients with HeFH (study group; n=36). DNA damage was determined by alkaline comet assay. PCSK9 protein level was assessed by ELISA. </span><span>The levels of Lp(a) in human serum were quantitatively turbidimetrically assay.</span><span> </span><span>PCSK9i treatment was associated with lower DNA damage, Lp(a), PCSK9 and lipid profile than before treatment. However, 20 of 36 patients still had Lp(a) values above 125 nmol/L, and reduced Lp(a) did not correlate with reduced DNA damage. Reduced PCSK9 moderately (r=0.48) correlates with reduced DNA damage</span><span><span>. PCSK9i therapy reduces the level of DNA damage in HeFH patients, regardless of the type of inhibitor. The reduction in DNA damage is not related to the changes in lipid profile or Lp(a) induced by PCSK9i, but it is dependent on PCSK9 level.</span></span></p>
CLDF dataset derived from Ugarte et al.'s "NorthPeruLex - A Lexical Dataset of Small Language Families and Isolates from Northern Peru (forthcoming).
<p>Cite the source of the dataset as:</p> <blockquote> <p>Ugarte, Carlos and Blum, Frederic and Ingunza, Adriano and Gonzales, Rosa and Peña, Jaime. Forthcoming. NorthPeruLex - A Lexical Dataset of Small Language Families and Isolates from Northern Peru.</p> </blockquote>
Most bacterial gene families are biased toward specific chromosomal positions
<p><span>The arrangement of genes along bacterial chromosomes influences their expression through growth rate-dependent gene copy number changes during DNA replication. While translation and transcription genes often cluster near the origin of replication, the extent of positional biases across gene families remains unclear. We hypothesized that natural selection broadly favors specific chromosomal positions to optimize growth rate-dependent expression. Analyzing 910 bacterial species and proteomics data from <em>Escherichia coli</em> and <em>Bacillus subtilis</em>, we find that about two-thirds of bacterial gene families are positionally biased, mainly near the origin or terminus of replication, with the strongest natural selection in fast-growing species. Our findings reveal chromosomal positioning as a fundamental mechanism for coordinating gene expression with growth rate, highlighting evolutionary constraints on bacterial genome architecture.</span></p> <p> </p>
Metagenomes: gene function and family annotations
<p>Functional annotations of genes for all contigs in 1,782 metagenomes.</p> <p>Genes were annotated to three sources: (1) COGs, (2) Pfams, and (3) <em>de novo</em> families from reference sequences. These gene annotations are used to train and run PlasX.</p>
Datasets for "The Venturia inaequalis effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins "
<p>Datasets for preprint entitled "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi"</p> <p><strong>1) ViAnnotation.gff3</strong><br> Gene annotation of <em>Venturia inaequalis</em> MNH120 (<a href="https://genome.jgi.doe.gov/Venin1/Venin1.home.html">https://genome.jgi.doe.gov/Venin1/Venin1.home.html</a>) generated as part of the study "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi". </p> <p>Gene reannotation was performed to include genes that would have been missed in the previous annotation by Deng et al. (2017), especially those genes encoding putative effector proteins, which are difficult to predict. For this purpose, we used a three-step approach. In the first step, coding sequences (CDSs) from <em>V. inaequalis</em> isolate 05/172, which were predicted as part of a previous study by Passey et al. (2018) (<a href="https://journals.asm.org/doi/full/10.1128/MRA.01062-18">https://journals.asm.org/doi/full/10.1128/MRA.01062-18</a>), were downloaded from the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/">https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/</a>) and mapped to the MNH120 genome using GMAP v2021-02-22. In the second step, RNA-seq reads from one biological replicate representing each <em>in planta</em> time point of <em>Malus domestica</em> infection by <em>V. inaequalis </em>(12 hour post-inoculation [hpi], 24 hpi, 2 days post-inoculation [dpi], 3 dpi, 5 dpi, 7 dpi), as well as one time point representing growth of the fungus in culture, were mapped to the MNH120 genome using HISAT2 v2.2.1. Then, a genome-guided <em>de novo</em> transcriptome assembly was performed using Trinity v2.12.0 and likely CDSs were identified using Transdecoder v5.5.0 (<a href="https://github.com/TransDecoder/TransDecoder">https://github.com/TransDecoder/TransDecoder</a>) in conjunction with a minimum open frame (ORF) length of 50 amino acids. Finally, in the third step, all annotations were visualized in Geneious v9.05, together with the previous annotation from Deng et al. (2017), and a manual curation was performed to create a consensus prediction. Note: this reannotation was generated with the aim of identifying as many genes as possible, and as a result, it contains many spurious genes. </p> <p><strong>2) Protein_sequences_ViAnnotation.fasta</strong></p> <p><strong>3) ECs_Families_AlphaFold.zip</strong></p> <p>This dataset is made up of predicted protein tertiary structures representing the main member of each up-regulated <em>V. inaequalis</em> effector candidate family. Structures were predicted using Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>). In cases where the effector candidate had less than 30 proteins with amino acid sequence similarity in the NCBI database, a custom multiple sequence alignment (MSA) was generated and used as input for AlphaFold2. Here, mature protein sequences were used.</p> <p><strong>4) singletons_AlphaFold_OpenSourceCASP14.zip</strong></p> <p>This dataset set is made up of predicted protein tertiary structures representing up-regulated<em> V. inaequalis</em> singleton effector candidates. Structures were predicted using AlphaFold (<a href="https://github.com/deepmind/alphafold">https://github.com/deepmind/alphafold</a>) open source code v2.0.1 and v2.1.0, with pre-set casp14, max_template_date: 2020-05-14. Mature protein sequences were used as input. </p> <p><strong>5) ECs_Avrs_phytopathogens_AlphaFold.zip</strong></p> <p>Predicted tertiary structures of avirulence (Avr) proteins or candidate Avr proteins from other fungal pathogens included in the "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi" study. These structures were predicted using Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>). Mature protein sequences were used as input. </p> <p>If you have any questions about the datasets, please contact us.<br> Mercedes Rocafort: <a href="mailto:m.rocafort.ferrer@massey.ac.nz">m.rocafort.ferrer@massey.ac.nz</a><br> Carl Mesarich: <a href="mailto:c.mesarich@massey.ac.nz">c.mesarich@massey.ac.nz</a></p>
Genomic incongruence accompanies the evolution of flower symmetry in Eudicots: a case study in the poppy family (Papaveraceae, Ranunculales)
<p>Nuclear and plastid datasets and phylogenomic workflow associated to "Genomic Incongruence Accompanies the Evolution of Flower Symmetry in Eudicots: a case study in the poppy family (Papaveraceae, Ranunculales)", published in <em>Frontiers in Plant Science </em>15:1340056.<br>This compressed file (poppy_repo.zip) contains a markdown readme file (poppy_readme.md) describing the phylogenomic workflow followed, as well as two dataset folders (poppy_nuc and poppy_pl) divided into four (aln_nuc, gtr_nuc, sptr_nuc, and chrono_nuc) and three (aln_pl, sptr_pl, and chrono_pl) subfolders, respectively.<br>The nuclear folder (poppy_nuc) comprises shrunk and trimmed alignments (aln_nuc), ML gene trees (gtr_nuc), coalescent species trees (sptr_nuc), and a time tree (chrono_nuc).<br>The plastid folder (poppy_pl) comprises shrunk and trimmed alignments (aln_pl), a concatenated ML species tree (sptr_pl), and a time tree (chrono_pl).<br>The research article is available at https://www.frontiersin.org/journals/plant-science/articles/10.3389/fpls.2024.1340056 (doi: 10.3389/fpls.2024.1340056).</p>
CLDF dataset derived from Birchall et al.'s "A Combined Comparative and Phylogenetic Analysis of the Chapacuran Language Family" from 2016
<p>Cite the source of the dataset as:</p> <blockquote> <p>Birchall J, Dunn M, & Greenhill SJ. 2016. A Combined Comparative and Phylogenetic Analysis of the Chapacuran Language Family. International Journal of American Linguistics 82(3). 255–284.</p> </blockquote>
CLDF dataset derived from Blum et al.'s "A phylolinguistic classification of the Quechua language family" from 2023
<p>Cite the source of the dataset as:</p> <blockquote> <p>Blum, Frederic, Carlos Barrientos, Adriano Ingunza & Zoe Poirier. 2023. A phylolinguistic classification of the Quechua language family. INDIANA - Anthropological Studies on Latin America and the Caribbean 40(1). 29–-54. DOI: https://doi.org/10.18441/IND.V40I1.29-54.</p> </blockquote>
CLDF dataset derived from Chacon's "A revised proposal of Proto-Tukanoan consonants and Tukanoan family classification" from 2014
<p>Cite the source of the dataset as:</p> <blockquote> <p>Thiago Chacon. (2014). A revised proposal of Proto-Tukanoan consonants and Tukanoan family classification. Journal of American Linguistics 80.3, pp. 275–322. doi: https://doi.org/10.1086/676393</p> </blockquote>
CLDF dataset derived from Robinson and Holton's "Internal Classification of the Alor-Pantar Language Family" from 2012
<p>Cite the source of the dataset as:</p> <blockquote> <p>Robinson, Laura C. and Holton, Gary (2012): Internal Classification of the Alor-Pantar Language Family Using Computational Methods Applied to the Lexicon. Language Dynamics and Change 2.2. 123-149.</p> </blockquote>
Human Family With Sequence Similarity 83 Member B (FAM83B); A Target Enabling Package
<p>FAM83A-H are newly identified oncogenes characterised by a conserved DUF1669 domain. FAM83B can substitute for RAS to promote malignant transformation. Ablation of FAM83B or mutation of Lys230 inhibits malignant phenotypes, implicating FAM83B as potential therapeutic target. As part of this TEP, we solved the first crystal structures from the FAM83 family, including FAM83A and FAM83B. The structures of the DUF1669 domain reveal a phospholipase D-like fold lacking conservation of key catalytic residues. We deorphanise the FAM83 DUF1669 domain as a critical docking scaffold for binding of casein kinase 1 isoforms. Finally, using XChem fragment screening we report chemical fragments that bind to Lys230 in the central pocket of the DUF1669 and form starting points for potential drug development.</p>
Human Poly (ADP-ribose) Polymerase Family Member 14 (PARP14); A Target Enabling Package
<p>This work provides reagents to develop specific inhibitors of the macrodomains of PARP14, for potential use for cancer and/or inflammation. Targeting the macrodomains of PARP14 is an alternative targeting strategy to PARP catalytic domain inhibitors that may allow greater inhibitor selectivity, and an alternative cellular effect.</p> <p>This package includes protein purification protocols, crystal structures of the 2nd and 3rd macrodomain of PARP14 in complex with small molecule chemical starting points, <em>in vitro </em>assays to measure ligand binding to the macrodomains, as well as validation of a PARP14 antibody and reagents to generate PARP14 knock-out cell lines (CRISPR-Cas9).</p>
Edward FitzGerald Life and Letters – Analysis of family and friends
<p>These files form part of an archive of research material relating to the Life and Letters of Edward FitzGerald. The data have been compiled by independent researchers W H (Bill) Martin and Sandra Mason; their contact details are below. The files comprise a number of searchable listings of information contained in FitzGerald’s letters. The information formed input to a book on Edward FitzGerald which is referenced below.</p> <p>This section of the archive contains a database, <em>efgpeople</em>, which analyses the letters in terms of their comments on FitzGerald’s family and a wide range of his friends. It also shows the dating of the letters, the people to whom FitzGerald wrote, and his location at the time of writing. The letters are those contained in the collection published by A M Terhune and A B Terhune in 1980 – see reference below. The database contains all the letters published by the Terhunes, including some which have no references to family and friends. An explanatory README text file contains a table showing the fields included in the database and giving definitions of them and of the codings used where relevant. </p>
Dataset for genes and gene variants from familial gastroschisis
<p>The dataset consists of records from whole exome sequecing and bioinformatic analysis which includes genes and gene variants from a Mexican family with recurrence for gastroschisis (two affected half-sisters with gastroschisis, mother, and father of the proband).</p> <p>Release of this dataset was based on the Human Genome annotation, GRCh37/hg19.</p> <p>The full list of tables is described in the file READ ME and remain available in csv files.</p> <p> </p>
Frequency and Rank of Family Names in Peru
<p>Count of family names (surnames, last names) in Peru, from an approximately 7% sample of the adult population.</p> <p>In Peru, many people are registered as supporters of political parties, and their names are published by the <a href="https://aplicaciones007.jne.gob.pe/srop_publico/Consulta/PadronAfiliado">Registro de Organizaciones Políticas</a>. The lists include a DNI (national identity number) for each person to avoid duplicates. The 1,572,002 people on these lists (excluding the regional movements) represent around 7% of the adult population of Peru.</p> <p>Their maternal and paternal family names have been sorted and counted. Nearly all of the names have entries for both paternal and maternal names.</p> <p>These 3,142,561 family names represent 85,395 different names, most of which are infrequent. The file has been limited to names that occur ten or more times in the sample, which is 12,139 unique names (3,021,655 names, more than 96% of the total).</p> <p>Each row in the file contains the rank, a percentage of that name in the entire set of 3,142,561 names, a count of the times the name occurs in the sample, and the name. </p> <p>There are some names (around 800) in this file that contain a space. In most cases, these are names like "GARCIA DE RUIZ", where RUIZ is the name of the woman's husband. There are also cases where the name is like "DE LA CRUZ", which is a complete family name. No attempt has been made to remove the part of names which refer to the husband's name, this could be considered for a later version. </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.