Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
20,966
datasets available to search
ShareScore release 0.9.0
Dataset results
20,966 results for “family”
Ecological Momentary Assessment of Family Forest Owners in New England 2016
Family forest owners (FFOs) across the U.S., and particularly in New England, are critically important to the health and future of the nation’s forests. It is important to understand the behavior of FFOs because they in the United States collectively own more forested land than the federal government or any other type of owner, and their actions and decisions will impact the public goods these forests provide. Typical studies of FFO behavior use self-reported survey data; participants are asked to recall past behavior or predict future behavior. These surveys are prone to bias, as it can be difficult to remember when things occur or to accurately predict what one will do in the future. Ecological momentary assessments, used commonly in medicine, are a fresh approach to measuring behavior by querying the subject in real-time. The PING project was designed to reduce this bias and provide a more accurate snapshot of how landowners engage with their land on a short-term basis. Participants in two experimental groups were invited to take part in a month-long survey, where the same questions were sent to them in the method of their choosing (text, via social media, or e-mail) once per week, asking about woodland engagement that week. Participants also took a pre- and post-survey to capture both demographics and feedback on the method. Over 61% of participants completed all 4 surveys and there was no statistically significant difference between the day a participant received their survey. Demographics were consistent with national statistics on woodland owners. A plurality of woodland owners in the study harvested timber for personal use and collected non-timber forest products. Finally, 86% of participants found the method of contact and number of questions reasonable, while 77% found the weekly contact reasonable. A plurality of respondents reported that their answers were typical of a given week, but that the survey question made them think more about their woods than t
Survey of Family Forest Owners Regarding Invasive Insects in the Connecticut River Watershed 2017
Forest insects have significant direct impacts on forest ecosystems; they are also generating new risks, uncertainties, and opportunities for forest landowners. Our research objective is to understand: (1) whether and how insect infestations are shifting land-use regimes in New England by altering human decision-making, (2) how these changes to human decisions may affect regional forest ecosystems and the provisioning of select ecosystem services, and (3) how subsequent changes to forest ecosystems, in turn, affect landowners. The dataset is a result of gathering information from a random sample of family forest owners (FFOs) in the Connecticut River Watershed; FFOs own approximately half of the private forestland in the region. We collected information on characteristics and perceptions of these FFOs as well as their intentions for their land if faced with the presence or threat of invasive forest insects. To better understand FFO intentions, contingent behavior questions populate the conjoint analysis format of the questionnaire. These are the foundational data that form our further efforts to simulate the impacts of insect dynamics and landowner behavior on regional forest ecosystems, including forest carbon stores, forest structure and composition, and timber yields. One of our overarching hypotheses is that forest land-use change in response to insects will have greater near-term ecological consequences than climate change or insects by themselves.
A genome-wide study of ruminants reveals two endogenous retrovirus families still active in goats
<p>Additional file - A genome-wide study of ruminants reveals two endogenous retrovirus families still active in goats </p>
Reprocessing of the dataset "Plasma Proteome Profiling Reveals the Effects of Weight Loss on the Apolipoprotein Family and Systemic Inflammation Status"
<p>Reprocessing of the MassIVE repository MSV000080596, originally generated to investigate the dynamic changes in the plasma proteomes of a cohort of individuals with obesity following weight loss and maintenance. The reprocessing included all samples from 52 individuals taken right after the weight-loss process and during the weight maintenance phase of the study (Weeks 0, 4, 13, 26, 39, and 52).</p> <p>We used the sequence database generated by ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) representing all populations from the 1000 Genomes Project (doi.org/10.5281/zenodo.10149277). For the search, SearchGUI version 4.3.1 and PeptideShaker version 3.0.0 were used with the X!Tandem and Tide search engines. The modification settings specified were carbamidomethylation of C as fixed and oxidation of M, deamidation of N and Q, Pyrrolidone of E and Q, and acetylation of protein N-terminus as variable modifications. The maximum peptide length was set to 40 amino acids and the precursor and fragment ion tolerances were set to 7 and 20 ppm, respectively. Resulting PSMs were processed as described in (doi.org/10.1021/acs.jproteome.3c00243) using Percolator version 3.5 provided with features based on peptide retention time (DeepLC version 1.1.2) and fragmentation predictors (MS2PIP version 3.9.0), and filtered at a 1% estimated FDR.</p> <p>The attached file contains all the peptide-spectrum matches identified at 1% FDR. The peptides have been annotated with transcripts, genes, and alleles using the ProHap Peptide Annotator v1.1 (<a href="https://github.com/ProGenNo/ProHap_PeptideAnnotator">https://github.com/ProGenNo/ProHap_PeptideAnnotator</a>).</p>
Dataset of Pedigree, genotypes, clinical and biochemical characteristics of families of Northeastern Mexico
<p>This dataset combines pedigree, genotypes, clinical and biochemical data of 37 families of Northeastern Mexico. Primary reference is the article:</p> <p>Gallardo‑Blanco, H.L., Villarreal‑Perez, J.Z., Cerda‑Flores, R.M., Figueroa, A., Sanchez‑Dominguez, C.N., Gutierrez‑Valverde, J.M. ... Martinez‑Garza, L.E. (2017). Genetic variants in KCNJ11, TCF7L2 and HNF4A are associated with type 2 diabetes, BMI and dyslipidemia in families of Northeastern Mexico: A pilot study. Experimental and Therapeutic Medicine, 13, 523-529. https://doi.org/10.3892/etm.2016.3990</p> <p><strong>If you use these data please cite the corresponding manuscript, which can be downloaded here:</strong></p> <p>https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5348709/</p> <p>https://www.spandidos-publications.com/10.3892/etm.2016.3990</p> <p>This dataset contains genotypes for the following SNPs:</p> <p>rs2986742</p> <p>rs4846051</p> <p>rs1801131</p> <p>rs1801133</p> <p>rs6541030</p> <p>rs12130799</p> <p>rs11208654</p> <p>rs1137100</p> <p>rs12405556</p> <p>rs3118378</p> <p>rs3737576</p> <p>rs10923931</p> <p>rs7554936</p> <p>rs3737787</p> <p>rs2516839</p> <p>rs1040404</p> <p>rs4670767</p> <p>rs7578597</p> <p>rs13400937</p> <p>rs10496971</p> <p>rs2627037</p> <p>rs1801262</p> <p>rs1569175</p> <p>rs2975760</p> <p>rs3792267</p> <p>rs10510228</p> <p>rs1801282</p> <p>rs3856806</p> <p>rs4955316</p> <p>rs9809104</p> <p>rs4607103</p> <p>rs6548616</p> <p>rs734873</p> <p>rs5400</p> <p>rs2030763</p> <p>rs4402960</p> <p>rs1513181</p> <p>rs9291090</p> <p>rs10010131</p> <p>rs10007810</p> <p>rs385194</p> <p>rs1799883</p> <p>rs2504853</p> <p>rs7754840</p> <p>rs7745461</p> <p>rs1800750</p> <p>rs1800629</p> <p>rs361525</p> <p>rs12200998</p> <p>rs2397060</p> <p>rs192655</p> <p>rs1044498</p> <p>rs4463276</p> <p>rs731257</p> <p>rs864745</p> <p>rs32314</p> <p>rs2330442</p> <p>rs4717865</p> <p>rs3173798</p> <p>rs10954737</p> <p>rs854555</p> <p>rs3917542</p> <p>rs662</p> <p>rs705308</p> <p>rs3943253</p> <p>rs751141</p> <p>rs1471939</p> <p>rs12544346</p> <p>rs13266634</p> <p>rs7844723</p> <p>rs2242103</p> <p>rs1408801</p> <p>rs10811661</p> <p>rs10511828</p> <p>rs12779790</p> <p>rs3793791</p> <p>rs4746136</p> <p>rs1111875</p> <p>rs10885390</p> <p>rs11196175</p> <p>rs7903146</p> <p>rs10885406</p> <p>rs12255372</p> <p>rs290487</p> <p>rs4918842</p> <p>rs2237892</p> <p>rs10839880</p> <p>rs1837606</p> <p>rs5210</p> <p>rs5218</p> <p>rs5219</p> <p>rs2946788</p> <p>rs11227699</p> <p>rs7930460</p> <p>rs1800849</p> <p>rs1387153</p> <p>rs948028</p> <p>rs2270031</p> <p>rs2416791</p> <p>rs7961581</p> <p>rs2070586</p> <p>rs1503767</p> <p>rs2269793</p> <p>rs8050136</p> <p>rs818386</p> <p>rs2966849</p> <p>rs1879488</p> <p>rs757210</p> <p>rs2033111</p> <p>rs11652805</p> <p>rs10512572</p> <p>rs2125345</p> <p>rs12946618</p> <p>rs12946115</p> <p>rs12950541</p> <p>rs1885088</p> <p>rs3907047</p> <p>rs2071023</p> <p>rs2833479</p> <p>rs2833483</p> <p>rs2300386</p> <p>rs2835370</p> <p>rs1296819</p> <p>rs1892848</p> <p>rs4821004</p> <p> </p>
Diffraction images used to solve the structures published in the article "Exploration of Strategies for Mechanism-Based Inhibitor Design for Family GH99 endo-alpha-1,2-Mannanases."
<p>Raw diffraction images used for generating the structures published in the article "Exploration of Strategies for Mechanism-Based Inhibitor Design for Family GH99 endo-a-1,2-Mannanases" (available <a href="https://doi.org/10.1002/chem.201800435">here</a>). Full single-crystal datasets are published. The software used for the processing of each dataset is listed in their respective PDB entries.</p> <p> </p> <p>If you find this useful, please contact me at <a href="mailto:lukasz.sobala@hirszfeld.pl">lukasz.sobala@hirszfeld.pl</a>, I am just interested in how these data are used!</p>
Scholarly journals publishing articles by family and community physicians in Brazil, up to December 2018
<p>This is the dataset of manuscript titled "In which journals do family and community physicians in Brazil publish? The <em>Trajetórias MFC</em> project". There are two spreadsheets: the dataset proper and the data dictionary. See the manuscript for background.</p> <p>All spreadsheets are in the CSV (comma-separated values) format, delimited with semicolons and encoded in UTF-8 with the byte-order mark (BOM). The spreadsheets can be opened with desktop or Web application software (LibreOffice Calc, Microsoft Excel, Google Sheets) or with statistical software such as R.</p> <p>A <a href="https://zenodo.org/record/3905255">previous version</a> of this dataset was used in a <a href="https://doi.org/10.1101/19005744">preprint</a>. This version should be cited by an upcoming article.</p> <p>See also the <a href="https://doi.org/10.1136/fmch-2020-000321">article</a>, <a href="https://doi.org/10.5281/zenodo.3376310">dataset</a> and <a href="https://doi.org/10.5281/zenodo.3381576">supplementary table</a> for an earlier milestone, about the postgraduate education of family and community physicians in Brazil.</p>
Diffraction images used to solve the structures published in the article "Contribution of Shape and Charge to the Inhibition of a Family GH99 endo-α-1,2-Mannanase"
<p>Raw diffraction images used for generating the structures published in the article "Contribution of Shape and Charge to the Inhibition of a Family GH99 endo-α-1,2-Mannanase" (available <a href="https://doi.org/10.1021/jacs.6b10075">here</a>). Full single-crystal datasets are published. The software used for the processing of each dataset is listed in their respective PDB entries.</p> <p> </p> <p>If you find this useful, please contact me at <a href="mailto:lukasz.sobala@hirszfeld.pl">lukasz.sobala@hirszfeld.pl</a>, I am just interested in how these data are used!</p>
Diffraction images used to solve the structures published in the article "A Family of Dual-Activity Glycosyltransferase-Phosphorylases Mediates Mannogen Turnover and Virulence in Leishmania Parasites"
<p>Raw diffraction images used for generating the structures published in the article A Family of Dual-Activity Glycosyltransferase-Phosphorylases Mediates Mannogen Turnover and Virulence in Leishmania Parasites" (available <a href="https://doi.org/10.1016/j.chom.2019.08.009">here</a>). The software used for the processing of each dataset is listed in their respective PDB entries.</p> <p> </p> <p>If you find this useful, please contact me at <a href="mailto:lukasz.sobala@hirszfeld.pl">lukasz.sobala@hirszfeld.pl</a>, I am just interested in how these data are used!</p>
Diffraction images used to solve the structures published in the article "From 1,4-Disaccharide to 1,3-Glycosyl Carbasugar: Synthesis of a Bespoke Inhibitor of Family GH99 Endo-α-mannosidase"
<p>Raw diffraction images used for generating the structures published in the article "From 1,4-Disaccharide to 1,3-Glycosyl Carbasugar: Synthesis of a Bespoke Inhibitor of Family GH99 Endo-α-mannosidase" (available <a href="https://doi.org/10.1021/acs.orglett.8b03260">here</a>). Full single-crystal datasets, including images that were not used in the final analyses, are published. The software used for the processing of each dataset is listed in their respective PDB entries. An additional 720 degree dataset is provided, which has been collected from the same crystal as PDB 6HMH. This dataset has not been used to solve the structure presented in the paper. It works very well as an example of sulfur SAD phasing.</p> <p> </p> <p>If you find this useful, please contact me at <a href="mailto:lukasz.sobala@hirszfeld.pl">lukasz.sobala@hirszfeld.pl</a>, I am just interested in how these data are used!</p>
Alignments used in "The evolution of the phenylpropanoid pathway entailed pronounced radiations and divergences of enzyme families"
<p>Alignments used in de Vries et al. (2021) "The evolution of the phenylpropanoid pathway entailed pronounced radiations and divergences of enzyme families" published as</p> <p>(1) a pre-print: https://doi.org/10.1101/2021.05.27.445924</p> <p>(2) in Plant Journal (in press)</p>
Supplementary data for draft genome of a member of the ascomycotal fungal genus Pseudopithomyces (family Didymosphaeriaceae)
<p><span lang="EN-US">We update our previous draft genome of a member of genus <em>Pseudopithomyces</em> (previously annotated as <em>Pseudopithomyces maydicus</em> strain SBW1, now reannotated as <em>Pseudopithomyces sp</em>. strain SBW1. The new draft genome is based on a hybrid assembly utilising both ONT and Illumina data. The draft genome is comprised of 43 contigs with a total length of 39.65Mbp. We predict 13,669 protein coding gene models, of which 4241 (31%) were annotated to KEGG Orthology. Taxonomic assignment to <em>Pseudopithomyces sp.</em> was supported by comparative analysis of extracted ITS regions, mitochondrial DNA sequence and whole genome comparisons using <em>k</em>-mer sketches. </span></p> <p> </p> <p><span lang="EN-US">The following items of Additional Data Files are made available in this repository:</span></p> <p><span lang="EN-US">Additional Data File 1: contigs.fasta</span></p> <p><span lang="EN-US">FASTA file of entire assembly. </span><span lang="EN-US"> </span></p> <p> </p> <p><span lang="EN-US">Additional Data File 2: draft_genome.fasta</span></p> <p><span lang="EN-US">FASTA file of draft whole genome sequence.</span></p> <p> </p> <p><span lang="EN-US">Additional Data File 3: ITS_full.fasta</span></p> <p><span lang="EN-US">FASTA file containing full length ITS sequences from contig 23 and contig 42.</span></p> <p><span lang="EN-US"> </span></p> <p><span lang="EN-US">Additional Data File 4: 2NJ47W4U013-Alignment.txt</span></p> <p><span lang="EN-US">Text file containing BLASTN alignments for ITS region contained on contig 23</span></p> <p><span lang="EN-US"> </span></p> <p><span lang="EN-US">Additional Data File 5: 2NJXH2GE013-Alignment.txt</span></p> <p><span lang="EN-US">Text file containing BLASTN alignments for ITS region contained on contig 42</span></p> <p><span lang="EN-US"> </span></p> <p><span lang="EN-US">Additional Data File 6: 2NM79EUG016-Alignment.txt</span></p> <p><span lang="EN-US">Text file containing BLASTN alignments for the mitochondrial genome from contig 40 </span></p> <p><span lang="EN-US"> </span></p> <p><span lang="EN-US">Additional Data File 6: sourmash_bc10_hy_pm1_3.txt</span></p> <p><span lang="EN-US">Text file containing the MASH similarities of the draft genome compared to 18,883 fungal genomes.</span></p>
Research data for "Hutters in the Zamoyski Family Entail - history of an environmentally conditioned social group"
<p>Research data for "Hutters in the Zamoyski Family Entail - history of an environmentally conditioned social group" (v1_2024)</p>
Alpha-Galactosaminidase family GH114 protein from Fusarium solani: X-ray diffraction images
<p>This submission includes h5-files with diffraction images recorded using the Dectris EIGER X 16M detector at the DIAMOND beamline I04. The model of the crystal structure and associated information can be found in the Protein Data Bank entry 9EP6. The model has P 31 2 1 symmetry and three molecules per asymmetric unit. This is a case of crystal pathology – partial disorder. There is electron density for the fourth molecule which could be modelled with occupancy 1/2 and would overlap with a symmetry-related molecule.</p>
DeepOrchidSeries: A Sentinel-2 Dataset to inform convolutional SDMs with twelve-month Sentinel-2 image time-series, Orchid family
<p><strong>Deep Species Distribution Modelling from Sentinel-2 Image Time-series: a Global Scale Analysis on the Orchid Family</strong> </p> <ul> <li><strong><em>DeepOrchidSeries</em></strong> dataset gathers Sentinel-2 image time-series around geolocated orchid occurrences. Seasonal evolutions of the habitats are captured in the twelve-month RGB/IR time-series with 640x640m spatial resolution. It allows novel Species Distribution Models (SDMs) coupled with convolutional networks to take advantage of both spatial and temporal information.</li> <li>Our <strong>associated article</strong> is describing the modeling choices made to shape this ambitious dataset. It is submitted to <a href="https://www.frontiersin.org/research-topics/18336/plant-biodiversity-science-in-the-era-of-artificial-intelligence">https://www.frontiersin.org/research-topics/18336/plant-biodiversity-science-in-the-era-of-artificial-intelligence</a>. We believe such global data, methods and scripts are valuable to the conservation ecology community and especially deep-SDMs users. To our knowledge, no similar ready-to-use dataset is available. In the article, the dataset's temporal dimension is proven to significantly improve SDMs performances.</li> <li><strong><em>sen2patch</em></strong> is the gitlab project gathering the code to create such dataset. It is available at <a href="https://gitlab.inria.fr/jestopin/sen2patch">https://gitlab.inria.fr/jestopin/sen2patch</a>.</li> <li><strong><em>DeepOrchidSeries.csv</em></strong> contains all occurrences-level information. <ul> <li>We advice to load it with: <pre><code class="language-python">import pandas as pd df = pd.read_csv("path/to/DeepOrchidSeries.csv", sep=';') df.columns ['gbifid', 'canonical_name', 'decimallatitude', 'decimallongitude', 'speciesKey', 'cell_index', 'bot_country', 'bot_code', 'lvl2_code', 'continent_code']</code></pre> <ul> <li>'gbifid' is the occurrences GBIF ID</li> <li>'canonical_name', is the species canonical name</li> <li>'decimallatitude', 'decimallongitude' are the species coordinates in decimal degrees</li> <li>'speciesKey' is the species GBIF unique identifier</li> <li>'cell_index' is a unique cell ID in a 0.0025° lon/lat grid partitioning the Earth (used to stratify train/val/test set by geographic blocks)</li> <li>'bot_country', 'bot_code', 'lvl2_code', 'continent_code' are geographic subdivisions defined in <a href="https://github.com/tdwg/wgsrpd">https://github.com/tdwg/wgsrpd</a> (code and string for WGSRPD level 1, the botanical countries)</li> </ul> </li> </ul> </li> <li> <p>Initial <a href="https://www.gbif.org/">GBIF</a> query DOI is <a href="http://https://doi.org/10.15468/dl.4bijtu">https://doi.org/10.15468/dl.4bijtu</a> (26 August 2019).</p> </li> <li><strong><em>DeepOrchidSeries.tar</em></strong> file contains the satellite image time-series and is available at <a href="https://lab.plantnet.org/deeporchidseries/">https://lab.plantnet.org/deeporchidseries/</a> <ul> <li><em>.tar</em> archive measure 286 GB and extends to 432 GB once decompressed.</li> <li>Image time-series relative tree paths are constructed from the occurrences unique GBIF IDs.</li> <li>For a given occurence <em>gbifid</em>, matching patches are located in: <em>final_dataset_by_gbifid/gbifid[-2:]/gbifid[-4:-2]</em>, <em>i.e.</em> in a first folder named with the <em>gbifid</em> last two numbers and a subfolder with the previous two ones. Example: the time-series files matching occurrence 2236837714 are located at <em>final_dataset_by_gbifid/14/77/</em>. </li> <li>Image time-series are composed of twelve 16 bits RGB <em>.png</em> and twelve 16 bits IR <em>.png</em> files containing data identical to the original L1C products, no lossy compression was made. There are one RGB and one IR .png file per month.</li> <li>Patches from month MM/YYYY of occurrence <em>gbifid</em> are named<em> </em><em>RGB_YYYY_MM_gbifid_.png</em> and <em>IR0_YYYY_MM_gbifid_.png</em>.</li> </ul> </li> <li><em><strong>models.zip</strong></em> is the archive containing the four PyTorch models weights described in our article and<strong><em> </em></strong><em><strong>inception_env.py</strong></em> the used Inception V3 architecture. <em><strong>index.json</strong></em> contains the dictionnary linking the models class indexes from 0 to 14128 with our labels <em>speciesKey</em>: {"class_index":speciesKey}.</li> </ul> <p> </p> <ul> <li><strong>ACKNOWLEDGMENTS</strong>: We warmly thank Alexander Zizka et al. for providing us the geographically and taxonomically curated set of Orchids occurrences. This dataset contains modified Copernicus Sentinel data and Copernicus Service information (2018). Sentinel-2 MSI data used were available at no cost from ESA Sentinels Scientific Data Hub.</li> </ul>
Academic Family Tree Data Export
<p>The Academic Family Tree is a live, crowdsourced project, that documents academic mentoring relationships across many fields. Data are updated continually. This snapshot was taken on 2024-10-18.</p> <p>We welcome inquiries about this dataset and are interested in learning about any new results you uncover. Contact: <a href="mailto:davids@ohsu.edu?subject=Academic%20Family%20Tree%20%2F%20Zenodo%20dataset">davids@ohsu.edu</a>.</p> <p>This dataset contains key tables from the Academic Family Tree, including information on names/institutions of academic mentoring relationships and semi-automated links of authors to publications and US grants (NSF, NIH only). A subset of publication and grant links have been validated by human users. To save space, only unique identifiers are included for publications (PMID, DOI) and grants (federal project number), without other metadata (author, title, journal, principle investigator, etc). These identifiers should be adequate to link to other databases. Also note, author-publication links are broken into several separate files. These files should be concatenated into a single table to generate a complete dataset. Some additional information is available here: <a href="https://academictree.org/export.php">https://academictree.org/export.php</a>.</p> <p>These data are associated with Liénard, J.F., Achakulvisut, T., Acuna, D.E. <em>et al.</em> Intellectual synthesis in mentorship determines success in academic careers. <em>Nature Communications</em> <strong>9, </strong>4840 (2018) (<a href="https://doi.org/10.1038/s41467-018-07034-y">https://doi.org/10.1038/s41467-018-07034-y</a>). Please cite this publication in work that uses this dataset.</p> <p>Funded by NSF Award 1933675.</p>
Dataset: Systematics of the color-polymorphic spider genus Cybaeolus, with comments on the phylogeny of the family Hahniidae (Araneae)
<p>Phylogenetic analysis of the spiders of the genus Cybaeulus, with outgroups in the marronoid clade. Data from six DNA markers, analyzed with maximum likelihood and parsimony.</p> <p><br>PHYLOGENETIC ANALYSIS</p> <p>We obtained sequences from 26 samples of the three known species of Cybaeolus, and of five additional species of Hahniidae. To these, we added legacy sequences of Cybaeolus and of other genera of Hahniidae, as well as representatives of the remaining families in the marronoid clade. For the new sequences, the extraction and amplification of DNA was made in the Laboratory of Molecular Tools at Museo Argentino de Ciencias Naturales (MACN), from tissues preserved in absolute alcohol at -18ºC. We targeted the markers histone H3 (H3), cytochrome oxidase subunit I (CO1), 28S ribosomal RNA (28S) and 16S ribosomal RNA (16S), previously used to estimate relationships of marronoid spiders (Wheeler et al., 2017). Details of extraction, primers and PCR protocols are the same as in Magalhaes & Ramírez (2022). Sequencing was outsourced to Macrogen Inc., South Korea. The resulting chromatograms were analyzed individually to detect contaminated sequences or ambiguous portions. In addition to these sequences obtained in the laboratory, we combined our data with additional sequences from previous work (Wheeler et al., 2017; Rivera-Quiroz et al., 2020), using the markers mentioned above plus 12S ribosomal RNA (12S) and 18S ribosomal RNA (18S). For the CO1 marker, additional sequences obtained by the Arachnology Division at MACN and deposited in the BOLDSYSTEMS platform (https://www.boldsystems.org/) were also used. Sequences were aligned with MAFFT Online v.7.463 (Katoh & Standley, 2013), using the L-INS-I algorithm. See Table 1 for list of vouchers and sequence identifiers.</p> <p>Maximum likelihood<br>For the maximum likelihood analyses we used the program IQ-TREE 2.2.0 (Minh et al., 2020), partitioning the data by marker, and selecting the best combination of partitions and evolution models by Bayesian information criterion (best fitting models were TPM2+I+G4 for H3, GTR+F+I+G4 for 18S, GTR+F+I+G4 for 16S and 12S together, GTR+F+I+G4 for CO1, and GTR+F+I+G4 for 28S). Since the relationships of outgroup taxa in the resulting trees were slightly different to that found in recent phylogenomic studies, we used the study of Gorneau et al. (2023) based on ultraconserved elements as a backbone topology to constrain our tree search, considering only the taxa in common with our analysis (see supplementary Fig. S1); this means that all the rest of the taxa are free to move anywhere during tree search. Support for groups (branches) was estimated by 1000 cycles of ultrafast bootstrapping. Ten independent runs were performed; of those, six converged into nearly identical log likelihood values (-57417.7725 to -57417.9604) and identical topologies; the tree with top-ranking log likelihood is presented in Results, after collapsing branches with bootstrap below 0.5. To estimate the support of an alternative topology with Cybaeolus as sister to the rest of the hahniids, we used TNT 1.6 (Goloboff & Morales, 2023) to modify the optimal tree placing Cybaeolus in such position, and asked for the frequency of the branch of interest (all hahniids except Cybaeolus) in the 1000 bootstrapped trees previously saved by IQTREE.<br>Ancestral character states for the arrangement of spinnerets (grouped; separated in a transversal line) were estimated by maximum likelihood on the optimal tree, using the R packages phytools and ape, under the models ER and ARD, and the best fitting model selected by the Akaike information criterion. </p> <p>Parsimony<br>For the parsimony analyses we used TNT 1.6. For the equal weights analysis, a heuristic search was made using a driven search with the default parameters of the “new technologies”, aiming for 10 independent hits to minimum length. The resulting trees were then submitted to an additional round of tree-bisection reconnection (TBR) branch swapping. These results were compared to a simpler search strategy of 300 random addition sequences, each followed by TBR, which produced 20 hits to minimal length. As both strategies reached the same trees with multiple independent hits, it is likely that the optimal trees were found. Finally, the strict consensus of all the optimal trees was obtained, and on this consensus the support values were calculated by means of 1000 bootstrap pseudoreplicates. </p>
Dataset and R script for the analysis in the article "Food waste between environmental education, peers, and family influence. Insights from primary school students in Northern Italy", Journal of Cleaner Production
<p>We hereby publish the dataset (with metadata) and the R script (R Core team, 2018) used for implementing the analysis presented in the paper "Food waste between environmental education, peers, and family influence. Insights from primary school students in Northern Italy", <em>Journal of Cleaner Production </em>(Piras et al., 2023). The dataset is provided in csv format with semicolons as separators and "NA" for missing data. The dataset includes all the variables used in at least one of the models presented in the paper, either in the main text or in the Supplementary Material. Other variables gathered by means of the questionnaires included as Supplementary Material of the paper have been removed. The dataset includes inputted values for missing data on independent variables. These were inputted using two approaches: last observation carried forward (LOCF) - preferred when possible - and last observation carried backward (LOCB). The metadata are presented as a PDF file.</p>
PFAM Protein Families Dataset for Machine Learning
<p>A cleaned dataset of protein sequences and protein families for classification. The dataset is exported from PFAM as of June 2023 and curated to achieve the following characteristics:</p> <ul> <li>only protein families included with >=100 sequences</li> <li>families with >2000 sequences are truncated and only represented by 2000 sequences (chosen randomly)</li> <li>only proteins with sequence lengths between 100 and 1000</li> <li>amino acid sequences are form PDB; chains are concatenated only if not similar</li> </ul> <p>The dataset is not balanced, numbers of sequences per family in PFAM and in in dataset are:</p> <pre><code>families: 62, sequences: 46872 total (in PFAM) -> included (in dataset) Number in family ALLERGEN: 122 -> 122 Number in family APOPTOSIS: 381 -> 381 Number in family BIOSYNTHETIC PROTEIN: 346 -> 346 Number in family BIOTIN BINDING PROTEIN: 165 -> 165 Number in family BLOOD CLOTTING: 138 -> 138 Number in family CALCIUM BINDING PROTEIN: 135 -> 135 Number in family CELL ADHESION: 1116 -> 1116 Number in family CELL CYCLE: 511 -> 511 Number in family CHAPERONE: 964 -> 964 Number in family CONTRACTILE PROTEIN: 158 -> 158 Number in family CYTOKINE: 191 -> 191 Number in family DE NOVO PROTEIN: 253 -> 253 Number in family DNA BINDING PROTEIN: 1008 -> 1008 Number in family ELECTRON TRANSPORT: 841 -> 841 Number in family FLUORESCENT PROTEIN: 348 -> 348 Number in family GENE REGULATION: 607 -> 607 Number in family HORMONE: 272 -> 272 Number in family HORMONE GROWTH FACTOR: 159 -> 159 Number in family HORMONE RECEPTOR: 121 -> 121 Number in family HYDROLASE: 19551 -> 2000 Number in family HYDROLASE ANTIBIOTIC: 120 -> 120 Number in family HYDROLASE HYDROLASE INHIBITOR: 2890 -> 2000 Number in family HYDROLASE INHIBITOR: 315 -> 315 Number in family IMMUNE SYSTEM: 3333 -> 2000 Number in family IMMUNOGLOBULIN: 155 -> 155 Number in family ISOMERASE: 2457 -> 2000 Number in family ISOMERASE ISOMERASE INHIBITOR: 139 -> 139 Number in family LECTIN: 139 -> 139 Number in family LIGASE: 1780 -> 1780 Number in family LIGASE LIGASE INHIBITOR: 163 -> 163 Number in family LIPID BINDING PROTEIN: 421 -> 421 Number in family LIPID TRANSPORT: 115 -> 115 Number in family LUMINESCENT PROTEIN: 221 -> 221 Number in family LYASE: 4150 -> 2000 Number in family LYASE LYASE INHIBITOR: 298 -> 298 Number in family MEMBRANE PROTEIN: 1338 -> 1338 Number in family METAL BINDING PROTEIN: 951 -> 951 Number in family METAL TRANSPORT: 409 -> 409 Number in family MOTOR PROTEIN: 195 -> 195 Number in family OXIDOREDUCTASE: 11531 -> 2000 Number in family OXIDOREDUCTASE OXIDOREDUCTASE INHIBITOR: 766 -> 766 Number in family OXYGEN STORAGE: 127 -> 127 Number in family OXYGEN STORAGE TRANSPORT: 260 -> 260 Number in family OXYGEN TRANSPORT: 414 -> 414 Number in family PHOTOSYNTHESIS: 173 -> 173 Number in family PLANT PROTEIN: 255 -> 255 Number in family PROTEIN BINDING: 1613 -> 1613 Number in family PROTEIN TRANSPORT: 693 -> 693 Number in family RECEPTOR: 108 -> 108 Number in family REPLICATION: 161 -> 161 Number in family RNA BINDING PROTEIN: 546 -> 546 Number in family SIGNALING PROTEIN: 2312 -> 2000 Number in family STRUCTURAL PROTEIN: 869 -> 869 Number in family SUGAR BINDING PROTEIN: 1250 -> 1250 Number in family TOXIN: 546 -> 546 Number in family TRANSCRIPTION REGULATION: 3283 -> 2000 Number in family TRANSFERASE: 14724 -> 2000 Number in family TRANSFERASE INHIBITOR: 126 -> 126 Number in family TRANSFERASE TRANSFERASE INHIBITOR: 2465 -> 2000 Number in family TRANSLATION: 370 -> 370 Number in family TRANSPORT PROTEIN: 2782 -> 2000 Number in family VIRAL PROTEIN: 2150 -> 2000</code></pre> <p>Files:</p> <ul> <li>families.csv: list of protein families with frequencies</li> <li>pfam_46872x62.csv: full dataset with amino acid sequences as string (one-letter code)</li> <li>pfam-trn-xy.csv: training dataset with amino acid sequences as tokens (1..25) and padded to a common length of 1000 with padding token 0:</li> </ul> <pre><code> Amino acid | Token | Description -------------------------------- C | 1 | Cysteine S | 2 | Serine T | 3 | Threonine A | 4 | Alanine G | 5 | Glycine P | 6 | Proline D | 7 | Aspartic acid E | 8 | Glutamic acid Q | 9 | Glutamine N | 10 | Asparagine H | 11 | Histidine R | 12 | Arginine K | 13 | Lysine M | 14 | Methionine I | 15 | Isoleucine L | 16 | Leucine V | 17 | Valine W | 18 | Tryptophan Y | 19 | Tyrosine F | 20 | Phenylalanine B | 21 | Aspartic acid or Asparagine Z | 22 | Glutamic acid or Glutamine J | 23 | Leucine or Isoleucine U | 24 | Selenocysteine X | 25 | Unknown amino acid . | 0 | padding token</code></pre> <p> </p> <ul> <li>pfam-trn-labels.csv: plain-text labels for training data</li> <li>pfam-tst-xy.csv</li> <li>pfam-tst-labels.csv: test data</li> <li>pfam-balanced-trn-xy.csv</li> <li>pfam-balanced-trn-labels.csv:</li> <li>pfam-balanced-tst-xy.csv</li> <li>pfam-balanced-tst-labels.csv: balanced datasets, created by oversampling.</li> </ul>
Gene family expansions underpin context-dependency of the oldest mycorrhizal symbiosis
<p>This Zenodo archive is associated with the manuscript:</p> <p>Hernandez, D.J., Pohlmann, G.B., Afkhami, M.E. (2025) <span>Gene family expansions provide molecular flexibility required for context-dependent species interactions.</span> Ecology Letters.</p> <p>Abstract:</p> <p>As environments worldwide change at unprecedented rates during the Anthropocene, understanding context-dependency – how species regulate interactions to match changing environments – is crucial. However, generalizable molecular mechanisms underpinning context-dependency remain elusive. Combining comparative genomics across 42 angiosperms with transcriptomics, genome-wide association mapping, and gene duplication origin analyses, we show for the first time that gene family expansions undergird context-dependent regulation of species interactions. Gene families expanded in mycorrhizal fungi-associating plants display up to 200% more context-dependent gene expression and double the genetic variation associated with mycorrhizal benefits to plant fitness. Moreover, we discover these gene family expansions arise primarily from tandem duplications with >2-times more tandem duplications genome-wide, indicating gene family expansions continuously supply genetic variation throughout plant evolution allowing fine-tuning of context-dependency in species interactions.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.