Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

20,966

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

20,966 results for “family”

Learn how ShareScore rates datasets ↗
zenodo44/100

CLDF dataset derived from Galucio et al.'s "Lexical Distances within the Tupian Linguistic family" from 2015

<p>Cite the source of the dataset as:</p> <blockquote> <p>Galucio, Ana Vilacy and Meira, Sérgio and Birchall, Joshua and Moore, Denny and Gabas Júnior, Nilson and Drude, Sebastian and Storto, Luciana and Picanço, Gessiane and Rodrigues, Carmen Reis. (2015). Genealogical relations and lexical distances within the Tupian linguistic family. Boletim do Museu Paraense Emílio Goeldi. Ciências Humanas, 10(2), 229-274. https://dx.doi.org/10.1590/1981-81222015000200004</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

GBIF map data latest examples: family Gadidae

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo44/100

Araneae families abundance in pitfall traps in sub project 7 in KiLi project

<p>Abundances of spiders, resolution at least family level, FER2 was not sampled.</p> <p>We sampled arthropod assemblages in disturbed and undisturbed vegetation types along an elevational gradient of 860&ndash;4550 m asl on the southern slopes of Mt. Kilimanjaro, Tanzania. On each site, ten pitfall traps were evenly spaced along two 50 m transects, with a distance of 10 m between individual traps and 20 m between transects. Pitfall traps were filled with 100&ndash;200 ml of a mixture of ethylene glycol and water (1:1 vol/vol) with a drop of liquid soap to break surface tension. Traps were exposed for 7 days each during two to five sampling events in both the dry and wet seasons between May 2011 and October 2012. As the number of individuals collected in ten traps was very high, we had to confine the sorting and subsequent analysis to sub-sets of at least three traps per sampling site and sampling event. Unfortunately, we had to find out later that the ethylen glycol procured locally was actually a mixture of ethylen glycol and 2-ethoxyethanol, which is a strong oxidizing chemical. Therefore, any sequencing of specimen caught in pitfall traps was impossible.</p> <p>Haas, Michael. 2014. Master thesis. The influence of elevation on community composition and trophic position of spiders. University of Marburg</p> <p>The KiLi project (2010-2018) is a German Science Foundation (DFG) funded research unit (DFG research unit FOR1246) that focuses on biodiversity and ecosystem processes along altitudinal and disturbance gradients on Mt. Kilimanjaro (Tanzania, Africa), capitalizing on its world-wide unique range of climatic and vegetation zones. The research unit comprises 2 central projects and 7 subprojects from various disciplines. On a total of 60 study sites in both natural and human-disturbed ecosystems biodiversity (e.g. plants, soil arthropods, ants, bees, frogs, lizards, bats, birds), related ecosystem processes (decomposition, seed dispersal, pollination, herbivory, predation), and biogeochemical processes and properties of ecosystems (climate, soil properties and nutrient status, regulation of water and carbon fluxes, trace gas emissions, primary productivity, functional diversity) are analyzed.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

CLDF dataset derived from Gerardi and Reichert's "The Tupí-Guaraní Language Family: A Phylogenetic Classification" from 2021

<p>Cite the source of the dataset as:</p> <blockquote> <p>Ferraz Gerardi, Fabrício and Reichert, Stanislav (2021) The Tupí-Guaraní Language Family: A Phylogenetic Classification. Diachronica 38(2). 151--188. DOI: https://doi.org/10.1075/dia.18032.fer.</p> </blockquote>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Gerational issues in linking family farming production, traditional food in diet, physical activity and obesity in Pacific Islands countries and territories: the case of the Melanesian population on Lifou Island

<p>In the Melanesian culture, traditional activities are organized around family farming, although the lifestyle transition taking place over the last several decades has led to imbalances in diet and physical activity, with both leading to obesity. The aim of this interdisciplinary study was to understand the links between family farming (produced, exchanged, sold, and consumed food), diet (focused on produced, hunted, and caught food), physical activity (sedentary, light, and moderate-to-vigorous physical activity) and obesity in Melanesian Lifou Island families (parents and children). Forty families, including 142 adults and children, completed individual food frequency questionnaires, wore tri-axial accelerometers for seven continuous days, and had weight and height measured with a bio-impedance device. Qualitative and quantitative interviews were conducted at the household level concerning family farming practices and sociodemographic variables. Multinomial regression analyses and logistic regression models were used to analyze the data. Results showed that family farming production brings a modest contribution to diet and active lifestyles for the family farmers of Lifou Island. The drivers for obesity in these tribal communities were linked to diet in the adults, whereas parental socioeconomic status and moderate-to-vigorous physical activity were the main factors associated to overweight and obesity in children. These differences in lifestyle behaviors within families suggest a transition in cultural practices at the intergenerational level. Future directions should consider seasonality and a more in-depth analysis of diet including macro- and micro- nutrients to acquire more accurate information on the intergenerational transition in cultural practices and its consequences on health outcomes in the Pacific region.</p>

opencc-bySep 2021View details →
zenodo44/100

Database of Planar and Three-Dimensional Periodic Orbits and Families Near the Moon

<p>The lunarPOdatabase.zip is the digital database accompanying the paper:<br> <br> C. Franz and R. P. Russell, &ldquo;Database of planar and three-dimensional periodic orbits and families near the Moon,&rdquo; The Journal of the Astronautical Sciences, DOI 10.1007/s40295-022-00361-9 (accepted Nov. 2022).</p> <p>Please see the paper for details, and cite the paper as appropriate. The database is accessible and permanently archived with the following DOI <a href="https://nam12.safelinks.protection.outlook.com/?url=https%3A%2F%2Fdoi.org%2F10.5281%2Fzenodo.6411980&amp;data=05%7C01%7C%7C1770e232817c4c99fc1b08dad0a0e3ad%7C31d7e2a5bdd8414e9e97bea998ebdfe1%7C0%7C0%7C638051686701542157%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&amp;sdata=Spb%2FFj5YL8QKaBoojr09MCCIdloAkzLiGKHoUo3uVGE%3D&amp;reserved=0">https://doi.org/10.5281/zenodo.6411980</a>.&nbsp; See accompanying license.txt and gpl-3.0.txt for license information, applying to all files included in the .zip distribution.</p> <p>The database contains over 13 million planar and three-dimensional solutions in the Earth-Moon circular restricted three body problem, grouped into 34,000 family and sub-family clusters. The database exists as human readable text files with periodic orbits organized by clusters and other dynamical characteristics.&nbsp; The database contains the clustered data, a README file describing the output format, an interactive GUI, and a simple MATLAB script as a basic interface with the database. The data are split into five files, one for each of the planar prograde, planar retrograde, axial prograde, axial retrograde, and x-z cases. The results (i.e. initial conditions and relevant dynamical parameters of each converged periodic orbit) are contained in a human-readable text file where each row is a new solution. The data are sorted by cluster and ordered inside the cluster to form a smooth curve. Summary files are included for both the grid search and the clustering for each run. The input parameters to the grid search software are also included with each case for reproducibility. File sizes range from approximately 1.1GB to 2.6GB, with a total uncompressed file size of 5.4GB and a total compressed file size of 1.2GB.</p> <p>It is emphasized that the GUI and other MATLAB interface files are only a preliminary capability to ease interaction with the database.&nbsp; They may not be stable under future releases of MATLAB. On the initial use of the GUI, we recommend to restrict the data to a single value of N (say N=1 or N=16), otherwise the number of solutions may overwhelm the system memory.&nbsp; If a user has difficulties using the GUI, the user is encouraged to use the MATLAB code interfaces or interface with the text files directly. The text files containing the database are the primary product provided here, with the GUI and test scripts provided as a courtesy to help ease the database&#39;s use.</p> <p>Please send questions to <a href="mailto:cfranz21@gmail.com">cfranz21@gmail.com</a> and/or <a href="mailto:ryan.russell@utexas.edu">ryan.russell@utexas.edu</a>.</p>

opengpl-2.0Nov 2022View details →
zenodo44/100

Dataset for 53 gene families from 16 eukaryotes

<p>NEXUS file representing Guigo et al.&#39;s (1996) dataset for 53 gene families from 16 eukaryotes, used in a number of gene tree reconciliation studies. Original data from&nbsp;Guigo et al., this NEXUS encoding by Roderic Page.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Vertebrate gene family trees

<p>This is a dataset of phylogenies for 118 vertebrate gene families, used in two papers published by James Cotton and Roderic Page in 2002. The folder named &quot;genes.zip&quot; contains FASTA and&nbsp;NEXUS sequence files, and Newick-format&nbsp;tree files for each gene family. There is a PDF file &quot;suppl.pdf&quot; listing the names of each family. The file &quot;final_dataset.gml&quot; contains a graph of the taxonomic overlap in the gene trees. All 118 gene trees have been combined into the file &quot;final_dataset.gtr&quot; which is a NEXUS file with custom blocks recognised by GeneTree.</p> <table> <tbody> <tr> <th>HOVERGEN FAMILY CODE</th> <th>GENE FAMILY NAME</th> </tr> <tr> <td>FAM000030A</td> <td>wnt 5</td> </tr> <tr> <td>FAM000030B</td> <td>wnt 7</td> </tr> <tr> <td>FAM000030C</td> <td>wnt 11</td> </tr> <tr> <td>FAM000030D</td> <td>wnt/int 1</td> </tr> <tr> <td>FAM000030E</td> <td>wnt 4</td> </tr> <tr> <td>FAM000030F</td> <td>wnt 10/12</td> </tr> <tr> <td>FAM000030G</td> <td>wnt 3</td> </tr> <tr> <td>FAM000030H</td> <td>wnt 2</td> </tr> <tr> <td>FAM000030I</td> <td>wnt 8</td> </tr> <tr> <td>FAM000105</td> <td>rhodopsin</td> </tr> <tr> <td>FAM000214</td> <td>beta-B globin</td> </tr> <tr> <td>FAM000033</td> <td>protamine 1</td> </tr> <tr> <td>FAM000370</td> <td>PrP a prion-protein</td> </tr> <tr> <td>FAM000014</td> <td>growth hormone</td> </tr> <tr> <td>FAM000556</td> <td>Rag-1 recombination activation gene</td> </tr> <tr> <td>FAM001493</td> <td>c-mos proto-oncogene</td> </tr> <tr> <td>FAM001462</td> <td>tyrosine kinase / yes / fyn / src / lck</td> </tr> <tr> <td>FAM001041</td> <td>metallothionein</td> </tr> <tr> <td>FAM000364</td> <td>Ldh-2 lactate dehydrogenase-B (EC</td> </tr> <tr> <td>FAM000215</td> <td>alpha globin</td> </tr> <tr> <td>FAM000016</td> <td>placental lactogen - prolactin</td> </tr> <tr> <td>FAM000008</td> <td>insulin</td> </tr> <tr> <td>FAM000824</td> <td>phosphoglycerate kinase</td> </tr> <tr> <td>FAM000550</td> <td>neurotrophin-4 (NT-4)</td> </tr> <tr> <td>FAM001478</td> <td>tyrosine kinase receptor, c-fms oncogene</td> </tr> <tr> <td>FAM000173A</td> <td>guanine nucleotide-binding protein</td> </tr> <tr> <td>FAM000173B</td> <td>transducin alpha</td> </tr> <tr> <td>FAM000502</td> <td>cytochrome P-450 aromatase</td> </tr> <tr> <td>FAM000192</td> <td>alpha-fetoprotein / serum albumin</td> </tr> <tr> <td>FAM000627</td> <td>neurone-specific enolase</td> </tr> <tr> <td>FAM001232</td> <td>preprotrypsin (ta)</td> </tr> <tr> <td>FAM000664</td> <td>complement component 3 (C3)</td> </tr> <tr> <td>FAM000175</td> <td>Ras</td> </tr> <tr> <td>FAM000248</td> <td>alpha B-crystallin</td> </tr> <tr> <td>FAM000006</td> <td>insulin-like growth factor II</td> </tr> <tr> <td>FAM000639</td> <td>transthyretin (prealbumin)</td> </tr> <tr> <td>FAM001303</td> <td>butylcholinesterase (BCHE)</td> </tr> <tr> <td>FAM000242</td> <td>connexin / gap junction protein</td> </tr> <tr> <td>FAM000058</td> <td>dopamine D1 receptor</td> </tr> <tr> <td>FAM000055</td> <td>beta-3-adrenergic receptor .</td> </tr> <tr> <td>FAM000131</td> <td>ATPase (Na+K+, H+K+)</td> </tr> <tr> <td>FAM003983</td> <td>preprogastrin</td> </tr> <tr> <td>FAM000353</td> <td>vasopressin</td> </tr> <tr> <td>FAM000152</td> <td>acetylcholine receptor</td> </tr> <tr> <td>FAM001327</td> <td>peripherin, desmin, vimentin, GFAP</td> </tr> <tr> <td>FAM003199</td> <td>(C57BL/6J)</td> </tr> <tr> <td>FAM002881</td> <td>cytochrome P-450, 17a-hydroxylase (CYP17)</td> </tr> <tr> <td>FAM002789</td> <td>(MUAHRB-1) Ah-receptor (Ah)</td> </tr> <tr> <td>FAM001461</td> <td>tropomyosin</td> </tr> <tr> <td>FAM000385</td> <td>Y3 peptide supply factor</td> </tr> <tr> <td>FAM000378</td> <td>pancreatic polypeptide, neuropeptide Y</td> </tr> <tr> <td>FAM001329</td> <td>cytokeratin</td> </tr> <tr> <td>FAM001619</td> <td>amelogenin (enamel-specific protein)</td> </tr> <tr> <td>FAM000286</td> <td>glucagon</td> </tr> <tr> <td>FAM000330</td> <td>tissue inhibitor of</td> </tr> <tr> <td>FAM000475</td> <td>lipophilin</td> </tr> <tr> <td>FAM001060</td> <td>ornithine carbamoyltransferase</td> </tr> <tr> <td>FAM001328</td> <td>neurofilament</td> </tr> <tr> <td>FAM001370</td> <td>peroxisome proliferator</td> </tr> <tr> <td>FAM000371</td> <td>somatostatin</td> </tr> <tr> <td>FAM001607</td> <td>Wilms tumor assocated protein (WT1)</td> </tr> <tr> <td>FAM000495</td> <td>aldolase A, B, C</td> </tr> <tr> <td>FAM001365</td> <td>liver receptor homologous protein (LRH-1)</td> </tr> <tr> <td>FAM000672</td> <td>ribosomal protein S4,</td> </tr> <tr> <td>FAM000135</td> <td>Na, K-ATPase beta-1 subunit</td> </tr> <tr> <td>FAM000271</td> <td>enkephalin : 1 2.</td> </tr> <tr> <td>FAM000799</td> <td>RING10</td> </tr> <tr> <td>FAM001664</td> <td>glutamate decarboxylase</td> </tr> <tr> <td>FAM000617</td> <td>creatine kinase</td> </tr> <tr> <td>FAM001337</td> <td>amyloid beta protein precursor</td> </tr> <tr> <td>FAM000274</td> <td>basic fibroblast growth factor (bFGF)</td> </tr> <tr> <td>FAM001108</td> <td>Sl-d mutant allele kit ligand (KL)</td> </tr> <tr> <td>FAM000504</td> <td>anion exchange protein 3</td> </tr> <tr> <td>FAM001239</td> <td>prothrombin</td> </tr> <tr> <td>FAM001360</td> <td>high mobility group proteins HMG1 and HMG2</td> </tr> <tr> <td>FAM000534</td> <td>glutamine synthetase</td> </tr> <tr> <td>FAM001053</td> <td>nucleoside diphosphate kinase</td> </tr> <tr> <td>FAM001390</td> <td>low density lipoprotein receptor LDLR</td> </tr> <tr> <td>FAM001464</td> <td>Cek6 receptor tyrosine kinase</td> </tr> <tr> <td>FAM000801</td> <td>manganese-containing superoxide</td> </tr> <tr> <td>FAM003946</td> <td>Six2 / Six1 mRNA</td> </tr> <tr> <td>FAM001606</td> <td>Ikaros binding protein (Ikaros)</td> </tr> <tr> <td>FAM000350</td> <td>SPARC protein</td> </tr> <tr> <td>FAM001339</td> <td>calcium-binding protein</td> </tr> <tr> <td>FAM000553</td> <td>pyruvate kinase</td> </tr> <tr> <td>FAM000604</td> <td>t complex polypeptide 1 (Tcp-1-a)</td> </tr> <tr> <td>FAM002988</td> <td>mSlo</td> </tr> <tr> <td>FAM001266</td> <td>transcription factor / hepatocyte nuclear factor</td> </tr> <tr> <td>FAM000453</td> <td>terminal deoxynucleotidyltransferase</td> </tr> <tr> <td>FAM001642</td> <td>transformation associated protein p53</td> </tr> <tr> <td>FAM000300A</td> <td>glucose-regulated protein 78 / HSP70 PART A</td> </tr> <tr> <td>FAM000300B</td> <td>glucose-regulated protein 78 / HSP70 PART B</td> </tr> <tr> <td>FAM000492A</td> <td>alpha actin etc</td> </tr> <tr> <td>FAM000492B</td> <td>beta actin etc</td> </tr> <tr> <td>FAM000649</td> <td>Myelin Basic Protein</td> </tr> <tr> <td>FAM000843</td> <td>ribosomal protein S4</td> </tr> <tr> <td>FAM000170</td> <td>atrial natriuretic protein</td> </tr> <tr> <td>FAM000871A</td> <td>tyrosinase</td> </tr> <tr> <td>FAM000871B</td> <td>tyrosinase related protein 1</td> </tr> <tr> <td>FAM001605</td> <td>ZFX put. transcription activator</td> </tr> <tr> <td>FAM001479</td> <td>fibroblast growth factor</td> </tr> <tr> <td>FAM000160</td> <td>pro-opiomelanocortin (POMC)</td> </tr> <tr> <td>FAM001644</td> <td>fibrinogen alpha subunit</td> </tr> <tr> <td>FAM000564</td> <td>SNAP-25</td> </tr> <tr> <td>FAM000266</td> <td>nitric oxide synthase</td> </tr> <tr> <td>FAM004159</td> <td>chondroitin-6 sulfotransferase</td> </tr> <tr> <td>FAM000800</td> <td>Lmp-2 (LMPq) proteasome subunit</td> </tr> <tr> <td>FAM001595</td> <td>factor B</td> </tr> <tr> <td>FAM001134</td> <td>zona pellucida (ZP)</td> </tr> <tr> <td>FAM001632</td> <td>stromelysin-3</td> </tr> <tr> <td>FAM001366</td> <td>steroid receptor (TR2-9)</td> </tr> <tr> <td>FAM000526</td> <td>c-ski protein</td> </tr> <tr> <td>FAM002463</td> <td>thymosin beta 4 peptide</td> </tr> <tr> <td>FAM000904</td> <td>sequence-specific DNA-binding protein (AP-2)</td> </tr> <tr> <td>FAM001465</td> <td>T-cell specific tyrosine kinase (ltk)</td> </tr> <tr> <td>FAM000567</td> <td>triosephosphate isomerase</td> </tr> <tr> <td>FAM006113</td> <td>DNA-dependent RNA polymerase III, large subunit</td> </tr> <tr> <td>FAM001733</td> <td>DNA-dependent RNA polymerase II</td> </tr> </tbody> </table>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Systematics of the avian family Alaudidae using multilocus and genomic data

<p>SNP datasets (gzipped VCF&nbsp;files) and tree output (newick-formatted trees in .trees files) associated with the publication&nbsp;<em>Systematics of the avian family Alaudidae using multilocus and genomic data</em> by Alstr&ouml;m et al. (2023).</p> <p>In the VCF files, samples have intermediary sequence names. Below is a key (sorted per name in VCF file) to taxon and ID. Further information can be found in Dataset SM1 in the original publication.</p> <table> <tbody> <tr> <td><strong>VCF_name</strong></td> <td><strong>Taxon</strong></td> <td><strong>ID</strong></td> </tr> <tr> <td>A_hamertoni_hamertoni_1982_3_71</td> <td>Alaemon hamertoni hamertoni</td> <td>NHMUK 1982.3.71</td> </tr> <tr> <td>A_hamertoni_tertia_1982_3_12</td> <td>Alaemon hamertoni tertia</td> <td>NHMUK 1982.3.12</td> </tr> <tr> <td>A_phoenicura_1949_25_4559</td> <td>Ammomanes phoenicura</td> <td>NHMUK 1949.25.4559</td> </tr> <tr> <td>Ala_alaudipes</td> <td>Alaemon alaudipes alaudipes</td> <td>DZUG U5653&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> </tr> <tr> <td>Alauda_arvensis_KZ</td> <td>Alauda arvensis dulcivox</td> <td>DZUG U0587&nbsp;&nbsp;</td> </tr> <tr> <td>Alaudala_seebohmi</td> <td>Alaudala cheleensis seebohmi</td> <td>DZUG U5316</td> </tr> <tr> <td>AmmocincU4892_S121</td> <td>Ammomanes cinctura</td> <td>DZUG U4892</td> </tr> <tr> <td>AmmodesU5745_S28</td> <td>Ammomanes deserti</td> <td>DZUG U5745</td> </tr> <tr> <td>AmmodesU5748_S106</td> <td>Galerida cristata magna</td> <td>DZUG U5748</td> </tr> <tr> <td>C_duponti_margaritae_1952_51_107</td> <td>Chersophilus duponti margaritae</td> <td>NHMUK 1952.51.107</td> </tr> <tr> <td>Cal_brachydactyla</td> <td>Calandrella brachydactyla rubiginosa</td> <td>DZUG U5655</td> </tr> <tr> <td>CalbarlU5662_S32</td> <td>Calendulauda erythroclamys patae</td> <td>DZUG U5662</td> </tr> <tr> <td>CalsabotaU2344_S33</td> <td>Calendulauda sabota suffusca</td> <td>DZUG U2344</td> </tr> <tr> <td>CheralboU5652_S52</td> <td>Chersomanes albofasciata</td> <td>DZUG U5652</td> </tr> <tr> <td>E_leucotis_BMNH_1952.25.36</td> <td>Eremopterix leucotis</td> <td>NHMUK 1952.25.36</td> </tr> <tr> <td>ErealpeU4610_S144</td> <td>Eremophila alpestris brandti</td> <td>DZUG U4610</td> </tr> <tr> <td>Eremalauda_dunni</td> <td>Eremalauda eremodites</td> <td>DZUG U4578</td> </tr> <tr> <td>Eremopterix_hova_FMNH449163</td> <td>Eremopterix hova</td> <td>FMNH 449163</td> </tr> <tr> <td>G_deva_1949_whi_1_8153</td> <td>Galerida deva</td> <td>NHMUK 1949.whi.1.8153</td> </tr> <tr> <td>G_malabarica_1949_whi_1_7738</td> <td>Galerida malabarica</td> <td>NHMUK 1949.whi.1.7738</td> </tr> <tr> <td>HetarchU2811_S34</td> <td>Heteromirafra archeri (&quot;sidamoensis&quot;)</td> <td>DZUG U2811</td> </tr> <tr> <td>Heteromirafra_ruddi</td> <td>Heteromirafra ruddi</td> <td>DZUG U5661&nbsp;</td> </tr> <tr> <td>Lullula_arborea_U544</td> <td>Lullula arborea</td> <td>UWBM 64680&nbsp;</td> </tr> <tr> <td>M_a_africana_BMNH_1927.5.26.5</td> <td>Corypha [Mirafra] africana africana</td> <td>NHMUK 1927.5.26.5</td> </tr> <tr> <td>M_africana_athi_1951_13_2495</td> <td>Corypha [Mirafra] africana athi</td> <td>NHMUK 1951.13.2495</td> </tr> <tr> <td>M_africana_athi_1968_48_2</td> <td>Corypha [Mirafra] africana athi</td> <td>NHMUK 1968.48.2</td> </tr> <tr> <td>M_albicauda_1916_12_1_840</td> <td>Mirafra albicauda&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> <td>NHMUK 1916.12.1.840</td> </tr> <tr> <td>M_ashi_1982_3_4</td> <td>Corypha [Mirafra] ashi</td> <td>NHMUK 1982.3.4</td> </tr> <tr> <td>M_collaris_BMNH_1923.8.7.2633</td> <td>Amirafra [Mirafra] collaris</td> <td>NHMUK 1923.8.7.2633</td> </tr> <tr> <td>M_collaris_BMNH_1923.8.7.2634</td> <td>Amirafra [Mirafra] collaris</td> <td>NHMUK 1923.8.7.2634</td> </tr> <tr> <td>M_cordofanica_1932_8_6_261</td> <td>Mirafra cordofanica</td> <td>NHMUK 1932.8.6.261</td> </tr> <tr> <td>M_cordofanica_1932_8_6_262</td> <td>Mirafra cordofanica</td> <td>NHMUK 1932.8.6.262</td> </tr> <tr> <td>M_gilletti_1982_3_10</td> <td>Calendulauda [Mirafra] gilletti arorihensis</td> <td>NHMUK 1982.3.10</td> </tr> <tr> <td>M_gilletti_1982_3_9</td> <td>Calendulauda [Mirafra] gilletti degodiensis</td> <td>NHMUK 1982.3.9</td> </tr> <tr> <td>M_hypermetra_BMNH_1912.12.23.317</td> <td>Corypha [Mirafra] hypermetra hypermetra</td> <td>NHMUK 1912.12.23.317</td> </tr> <tr> <td>M_rufa_lynesi_1922_12_8_1553</td> <td>Calendulauda [Mirafra] rufa rufa&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> <td>NHMUK 1922.12.8.1553</td> </tr> <tr> <td>M_rufa_nigriticola_1932_8_6_253</td> <td>Calendulauda [Mirafra] rufa nigriticola</td> <td>NHMUK 1932.8.6.253</td> </tr> <tr> <td>M_somalica_somalica_1919_10_6_27</td> <td>Corypha [Mirafra] somalica somalica</td> <td>NHMUK 1919.10.6.27</td> </tr> <tr> <td>MelcalaU0583_S124</td> <td>Melanocorypha calandra psammochroa</td> <td>DZUG U0583</td> </tr> <tr> <td>MelyeltU0580_S35</td> <td>Melanocorypha yeltoniensis</td> <td>DZUG U0580</td> </tr> <tr> <td>Mirafra_alb_ZMB_2000.9488</td> <td>Mirafra albicauda&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> <td>ZMB 2000.9488</td> </tr> <tr> <td>Mirafra_albi_ZMB_49.225</td> <td>Mirafra albicauda&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> <td>ZMB 49.225 (holotype)</td> </tr> <tr> <td>Mirafra_degodiensis</td> <td>Calendulauda [Mirafra] gilletti degodiensis</td> <td>DZUG U2808</td> </tr> <tr> <td>Mirafra_rufocinnamomea_U5657</td> <td>Amirafra [Mirafra] rufocinnamomea fischeri</td> <td>FMNH 484672 (= DZUG U5657)</td> </tr> <tr> <td>MirchenU5660_S37</td> <td>Mirafra cheniana</td> <td>DZUG U5660</td> </tr> <tr> <td>MirchenU5759_S36</td> <td>Mirafra cheniana</td> <td>DZUG U5759</td> </tr> <tr> <td>MirfascU5758_S60</td> <td>Corypha [Mirafra] fasciolata</td> <td>DZUG U5758&nbsp;&nbsp;</td> </tr> <tr> <td>Panurus</td> <td>Panurus biarmicus</td> <td>1ET92164|SAMN13107499</td> </tr> <tr> <td>Ramphocoris</td> <td>Ramphocoris clotbey</td> <td>CEFE Rhcl1 (= DZUG U5651)</td> </tr> <tr> <td>Spizocorys_fringillaris_U5659</td> <td>Spizocorys fringillaris</td> <td>DZUG U5659</td> </tr> <tr> <td>Spizocorys_obbiensis_1908_5_28_104</td> <td>Spizocorys obbiensis</td> <td>NHMUK 1908.5.28.104</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>For convenience, the same data sorted per taxon:</p> <table> <tbody> <tr> <td><strong>VCF_name</strong></td> <td><strong>Taxon</strong></td> <td><strong>ID</strong></td> </tr> <tr> <td>Ala_alaudipes</td> <td>Alaemon alaudipes alaudipes</td> <td>DZUG U5653&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> </tr> <tr> <td>A_hamertoni_hamertoni_1982_3_71</td> <td>Alaemon hamertoni hamertoni</td> <td>NHMUK 1982.3.71</td> </tr> <tr> <td>A_hamertoni_tertia_1982_3_12</td> <td>Alaemon hamertoni tertia</td> <td>NHMUK 1982.3.12</td> </tr> <tr> <td>Alauda_arvensis_KZ</td> <td>Alauda arvensis dulcivox</td> <td>DZUG U0587&nbsp;&nbsp;</td> </tr> <tr> <td>Alaudala_seebohmi</td> <td>Alaudala cheleensis seebohmi</td> <td>DZUG U5316</td> </tr> <tr> <td>M_collaris_BMNH_1923.8.7.2633</td> <td>Amirafra [Mirafra] collaris</td> <td>NHMUK 1923.8.7.2633</td> </tr> <tr> <td>M_collaris_BMNH_1923.8.7.2634</td> <td>Amirafra [Mirafra] collaris</td> <td>NHMUK 1923.8.7.2634</td> </tr> <tr> <td>Mirafra_rufocinnamomea_U5657</td> <td>Amirafra [Mirafra] rufocinnamomea fischeri</td> <td>FMNH 484672 (= DZUG U5657)</td> </tr> <tr> <td>AmmocincU4892_S121</td> <td>Ammomanes cinctura</td> <td>DZUG U4892</td> </tr> <tr> <td>AmmodesU5745_S28</td> <td>Ammomanes deserti</td> <td>DZUG U5745</td> </tr> <tr> <td>A_phoenicura_1949_25_4559</td> <td>Ammomanes phoenicura</td> <td>NHMUK 1949.25.4559</td> </tr> <tr> <td>Cal_brachydactyla</td> <td>Calandrella brachydactyla rubiginosa</td> <td>DZUG U5655</td> </tr> <tr> <td>M_gilletti_1982_3_10</td> <td>Calendulauda [Mirafra] gilletti arorihensis</td> <td>NHMUK 1982.3.10</td> </tr> <tr> <td>M_gilletti_1982_3_9</td> <td>Calendulauda [Mirafra] gilletti degodiensis</td> <td>NHMUK 1982.3.9</td> </tr> <tr> <td>Mirafra_degodiensis</td> <td>Calendulauda [Mirafra] gilletti degodiensis</td> <td>DZUG U2808</td> </tr> <tr> <td>M_rufa_nigriticola_1932_8_6_253</td> <td>Calendulauda [Mirafra] rufa nigriticola</td> <td>NHMUK 1932.8.6.253</td> </tr> <tr> <td>M_rufa_lynesi_1922_12_8_1553</td> <td>Calendulauda [Mirafra] rufa rufa&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> <td>NHMUK 1922.12.8.1553</td> </tr> <tr> <td>CalbarlU5662_S32</td> <td>Calendulauda erythroclamys patae</td> <td>DZUG U5662</td> </tr> <tr> <td>CalsabotaU2344_S33</td> <td>Calendulauda sabota suffusca</td> <td>DZUG U2344</td> </tr> <tr> <td>CheralboU5652_S52</td> <td>Chersomanes albofasciata</td> <td>DZUG U5652</td> </tr> <tr> <td>C_duponti_margaritae_1952_51_107</td> <td>Chersophilus duponti margaritae</td> <td>NHMUK 1952.51.107</td> </tr> <tr> <td>M_a_africana_BMNH_1927.5.26.5</td> <td>Corypha [Mirafra] africana africana</td> <td>NHMUK 1927.5.26.5</td> </tr> <tr> <td>M_africana_athi_1951_13_2495</td> <td>Corypha [Mirafra] africana athi</td> <td>NHMUK 1951.13.2495</td> </tr> <tr> <td>M_africana_athi_1968_48_2</td> <td>Corypha [Mirafra] africana athi</td> <td>NHMUK 1968.48.2</td> </tr> <tr> <td>M_ashi_1982_3_4</td> <td>Corypha [Mirafra] ashi</td> <td>NHMUK 1982.3.4</td> </tr> <tr> <td>MirfascU5758_S60</td> <td>Corypha [Mirafra] fasciolata</td> <td>DZUG U5758&nbsp;&nbsp;</td> </tr> <tr> <td>M_hypermetra_BMNH_1912.12.23.317</td> <td>Corypha [Mirafra] hypermetra hypermetra</td> <td>NHMUK 1912.12.23.317</td> </tr> <tr> <td>M_somalica_somalica_1919_10_6_27</td> <td>Corypha [Mirafra] somalica somalica</td> <td>NHMUK 1919.10.6.27</td> </tr> <tr> <td>Eremalauda_dunni</td> <td>Eremalauda eremodites</td> <td>DZUG U4578</td> </tr> <tr> <td>ErealpeU4610_S144</td> <td>Eremophila alpestris brandti</td> <td>DZUG U4610</td> </tr> <tr> <td>Eremopterix_hova_FMNH449163</td> <td>Eremopterix hova</td> <td>FMNH 449163</td> </tr> <tr> <td>E_leucotis_BMNH_1952.25.36</td> <td>Eremopterix leucotis</td> <td>NHMUK 1952.25.36</td> </tr> <tr> <td>AmmodesU5748_S106</td> <td>Galerida cristata magna</td> <td>DZUG U5748</td> </tr> <tr> <td>G_deva_1949_whi_1_8153</td> <td>Galerida deva</td> <td>NHMUK 1949.whi.1.8153</td> </tr> <tr> <td>G_malabarica_1949_whi_1_7738</td> <td>Galerida malabarica</td> <td>NHMUK 1949.whi.1.7738</td> </tr> <tr> <td>HetarchU2811_S34</td> <td>Heteromirafra archeri (&quot;sidamoensis&quot;)</td> <td>DZUG U2811</td> </tr> <tr> <td>Heteromirafra_ruddi</td> <td>Heteromirafra ruddi</td> <td>DZUG U5661&nbsp;</td> </tr> <tr> <td>Lullula_arborea_U544</td> <td>Lullula arborea</td> <td>UWBM 64680&nbsp;</td> </tr> <tr> <td>MelcalaU0583_S124</td> <td>Melanocorypha calandra psammochroa</td> <td>DZUG U0583</td> </tr> <tr> <td>MelyeltU0580_S35</td> <td>Melanocorypha yeltoniensis</td> <td>DZUG U0580</td> </tr> <tr> <td>M_albicauda_1916_12_1_840</td> <td>Mirafra albicauda&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> <td>NHMUK 1916.12.1.840</td> </tr> <tr> <td>Mirafra_alb_ZMB_2000.9488</td> <td>Mirafra albicauda&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> <td>ZMB 2000.9488</td> </tr> <tr> <td>Mirafra_albi_ZMB_49.225</td> <td>Mirafra albicauda&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> <td>ZMB 49.225 (holotype)</td> </tr> <tr> <td>MirchenU5660_S37</td> <td>Mirafra cheniana</td> <td>DZUG U5660</td> </tr> <tr> <td>MirchenU5759_S36</td> <td>Mirafra cheniana</td> <td>DZUG U5759</td> </tr> <tr> <td>M_cordofanica_1932_8_6_261</td> <td>Mirafra cordofanica</td> <td>NHMUK 1932.8.6.261</td> </tr> <tr> <td>M_cordofanica_1932_8_6_262</td> <td>Mirafra cordofanica</td> <td>NHMUK 1932.8.6.262</td> </tr> <tr> <td>Panurus</td> <td>Panurus biarmicus</td> <td>1ET92164|SAMN13107499</td> </tr> <tr> <td>Ramphocoris</td> <td>Ramphocoris clotbey</td> <td>CEFE Rhcl1 (= DZUG U5651)</td> </tr> <tr> <td>Spizocorys_fringillaris_U5659</td> <td>Spizocorys fringillaris</td> <td>DZUG U5659</td> </tr> <tr> <td>Spizocorys_obbiensis_1908_5_28_104</td> <td>Spizocorys obbiensis</td> <td>NHMUK 1908.5.28.104</td> </tr> </tbody> </table> <p>Questions can be directed to the corresponding authors.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Supplementary Materials to "Subgrouping in a `dialect continuum': A Bayesian phylogenetic analysis of the Mixtecan language family"

<p>SM0: metadata on the languages of the sample</p> <p>SM 1: custom word list</p> <p>SM2: prose explanation of cognate coding and IPA conversion</p> <p>SM3: annotated cognate sets</p> <p>SM4: nexus files of the broad and fine grained cognate coding</p> <p>SM5: NeighborNet visualization with coloring by Josserand (1983)&#39;s groupings and by groupings from our analysis</p> <p>SM6: BEAST2 xml files</p> <p>SM7: MCC trees from BEAST2 analysis</p> <p>SM8: DensiTree visualization and visualization of full MCC tree of best performing model</p>

opencc-by-4.0May 2022View details →
zenodo44/100

DNA loss model explains the evolution of the neuropeptide LWamide, APGWamide, APGW/AKH, RPCH, AKH, ACP, CRZ, and GnRH families

<p><strong>R1: Establishment and purification of neuropeptide sequences</strong></p> <p>The LW, APGW, RPCH, AKH, CRZ, and GnRH neuropeptide families were searched in the GenBank database using 10 keywords: the neuropeptide name, the precursor abbreviation, the full name of the precursor, the full name of the precursor with the word &ldquo;prepropeptide,&rdquo; and the combinations of these terms. The candidate sequences were downloaded in FASTA format using the appropriate commands in the GenBank database. The AKH neuropeptide family was classified according to the groups published in the literature, as well as the amino acid number and sequence. Furthermore, the ACP hybrid family was identified in the GenBank database using BLAST alignments.</p> <p><strong>C00: Neuropeptide Precursor. </strong>Eight folders were named with the initials of each neuropeptide family. The AKH family folder was the only one containing four subfolders. All of the folders contained the same type of files: three text files named after the neuropeptide initials and the obtained result. The files identified with the words &ldquo;<em>with codes</em>&rdquo; contained the sequences with the codes generated for this study, whereas the documents with the word &ldquo;<em>Full</em>&rdquo; contained the GenBank database search results obtained with the 10 aforementioned keywords. These files were located in a folder named &ldquo;<em>Fasta Keywords.</em>&rdquo; Each file contained the results from each respective keyword. The files with the words &ldquo;<em>selected EA</em>&rdquo; contained the sequences that were selected for evolutionary analyses.</p> <p><strong>C01: BLAST ACP</strong>. The text file named &ldquo;00 BLAST ACP&rdquo; contains the BLAST alignment results obtained from the NCBI database generated with the Adipokinetic Hormone/Corazonin-related peptide from the transcriptome of <em>Callinectes toxotes</em>. The file named &ldquo;01 ACP Selected&rdquo; contains the precursors selected for this study. All sequences were in FASTA format and contained the codes summarized in Supplementary Material 3 &ldquo;<em>Database Sequences.</em>&rdquo;</p> <p>The file named &ldquo;<em>02 ACP selected EA</em>&rdquo; contains the ACP precursors of other species, which were used for the evolutionary analyses of <em>C. toxotes</em> ACP. The PDF file titled &ldquo;<em>03 ACP ProP 1.0 Serv</em>&rdquo; contains the results of the proteolytic cleavage sites of the precursors indicated in the file named &ldquo;<em>02 ACP selected EA,</em>&rdquo; which were generated using the aforementioned software.</p> <p><strong>C02: BLAST VP.</strong> The folder contains the results of the BLAST alignment against the NCBI database, which were generated with the virtual peptide sequences reported by Martinez-Perez et al. (2007). This folder contains seven text files. The name of each file corresponds to the precursor and species in which it was identified. Moreover, the PDF document named &ldquo;<em>Virtual peptides ProP 1.0 Serv</em>&rdquo; contains the results of the proteolytic cleavage sites generated with the aforementioned software.</p> <p><strong>C03: Debugging sequences with software.</strong> This folder contains three subfolders containing the results obtained with each software used in this study for the detection of each of the neuropeptide sequences using the appropriate keywords.</p> <p>The folder named &ldquo;<em>BioDataToolKit</em>&rdquo; contains six subfolders with the abbreviated name of each neuropeptide. Additionally, there is a file containing the sequences downloaded from the GenBank database, as well as a Microsoft Excel file containing the details generated by the software. The name of each file corresponds to the keywords used for each search. The software used in this study can be found in the following repository: <a href="https://github.com/rduarte24/BiodataToolkit">https://github.com/rduarte24/BiodataToolkit</a>.</p> <p>The folder named &ldquo;<em>Pro1.0Server</em>&rdquo; was organized in the same way as the results derived for the &ldquo;<em>BioDataToolKit</em>&rdquo; for each neuropeptide family. However, each of the neuropeptide folders contained a file with the pertinent sequences whereas another file contained the endoproteolytic cleavage sites of the neuropeptide precursors obtained with the software.</p> <p>The folder named &ldquo;Proteios&rdquo; contains seven files. The file names indicate the precursor analyzed with the software and the identified sequences in FASTA format. The Proteios software is available in the following website: <a href="https://github.com/Martin-Munive/Proteios">https://github.com/Martin-Munive/Proteios</a>.</p> <p><strong>C04: Neuropeptide precursors for evolutionary analysis.</strong> Files with the sequences of the neuropeptide precursors used for the generation of the phylogenetic trees in Supplementary Materials 4 and 7. The name of each file corresponds to the name of each of the analyzed neuropeptides.</p> <p><strong>R2: Transcriptome BLAST</strong></p> <p>Microsoft Excel file containing the BLAST alignments conducted using the sequences of the AKH/CRZ-related peptide (ACP) from <em>C. toxotes</em> and Corazonin (CRZ) from <em>C. arcuatus</em>. The following information is summarized in the spreadsheets named <em>C. toxotes</em> and <em>C. arcuatus</em>: Column A, neuropeptide name; Column B, species name; Columns C&ndash;G, BLAST alignment results; Column H, GenBank protein accession number; Column I, precursor sequence.</p> <p><strong>R3: </strong><strong>Construction of neuropeptide database</strong></p> <p>Microsoft Excel file with information pertaining to the database and a detailed description of each of the neuropeptide precursors analyzed in this study. The Excel file contains seven spreadsheet tabs. Each of the tabs contains the following columns:</p> <p><strong>Neuropeptides.</strong> Column A, sequence numbering in descending order; Column B, neuropeptide name; Column C, identification code used in this study; Column D, accession number; Columns E&ndash;G, species taxonomy; Columns H&ndash;L, GenBank sequence description; Columns M&ndash;N, literature reference and link. <strong>Taxonomy.</strong> Taxonomic description of each of the examined species derived from the NCBI database. <strong>Sequences evolutionary anal</strong>. This tab contains the code developed for this work in Column C; the GenBank accession codes of each neuropeptide are summarized in Column D and species taxonomy details are summarized in Columns E y F. <strong>Table of differences.</strong> Column B shows the codes of identical sequences and Column C shows the code of the sequence selected for this study. <strong>Codes deleted. </strong>This tab contains the accession codes of the species and the species name but contains no details on the properties of the neuropeptide precursors. <strong>Sequences Paper</strong>. Neuropeptide sequences reported in previous studies that were later reported in the GenBank database. The sequences marked with asterisks have not been previously reported in public databases. The codes used in this study to designate the sequences are also included. <strong>Keywords. </strong>Keywords used to conduct the GenBank database searches to obtain the members of each neuropeptide family.</p> <p><strong>R4: <em>In silico</em> validation, alignments, and phylogenetic relationships</strong></p> <p>Generated phylogenetic trees and results obtained from individual runs for each of the neuropeptide families with the DNA-LM and Kalign parameters using the IQ-TREE software.</p> <p>The folder named &ldquo;<em>RUN</em>&rdquo; contains the &ldquo;<em>DNALM and kalign 2.0 default parameters</em>&rdquo; subfolder. Both folders contain 11 subfolders with the names of each of the neuropeptide families, as well as the results obtained with the IQ-TREE software. The folder named &ldquo;<em>Trees</em>&rdquo; contains the folder &ldquo;<em>DNALM and kalign 2.0 default parameters</em>&rdquo; containing the phylogenetic trees for each of the neuropeptide families, which were created with the Itol software.</p> <p><strong>R5: BLAST alignment of the virtual peptide precursors</strong></p> <p>Results of the BLAST alignment of the virtual peptides described by Martinez-Perez et al. (2007) with respect to the sequences in the GenBank database. The files follow the same nomenclature as in the folder named &ldquo;<em>Carpeta 02 BLAST VP</em><strong>&rdquo;</strong> in Repository 1.</p> <p><strong>R6: Alignment of neuropeptide precursors</strong></p> <p>&ldquo;<em>DNALM and Kalign 2.0 default parameter</em>&rdquo; folders. Each of these folders contains the alignments of the examined neuropeptide precursors from each family and each folder is named after the corresponding neuropeptide. The remaining files contain the alignments in ascending order in the evolutionary scale and are appropriately named after the corresponding neuropeptide. The file named &ldquo;<em>All Sequence FASTA</em>&rdquo; contains the sequences used in our study in FASTA format.</p> <p><strong>R7: Phylogenetic clustering of the precursors </strong></p> <p>&nbsp;&ldquo;<em>DNALM and Kalign 2.0 default parameter</em>&rdquo; folders. Both folders contain the phylogenetic tree clustering results from Supplementary Material 6, which were obtained using the DNA-LM y Kalign parameters and the IQ-TREE software. All analyses were conducted using the GUANE-1 supercomputer (Universidad Industrial de Santander). The phylogenetic clustering results of all of the precursors are contained in the folders with the respective precursor name. The folder also contains Figure 6, which was included in our main manuscript.</p> <p>Additionally, a folder entitled &quot;Orthofinder and Robinson-Foulds&quot; is included, which corresponds to the analyses carried out for: the Robinson-Foulds metric and the Orthofinder software.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

PL11 - Single family house - Czerniewice (Poland)

<p>Data files for building: PL11 - Single family house - Czerniewice (Poland)</p><p>Languages: Polish, English</p><p>These files are part of the public benchmark repository created as a part of the crossCert EU project.&nbsp;</p><p>This repository contains curated building data, certificate results and, where available, measured performance results. The repository is publicly available so that it can be used as a testbench for new Energy Performance Certificate (EPC) procedures.</p><p>The files are organised in the following folders&nbsp; (note that not all files are always provided):</p><ol><li>Main data&nbsp; and Results, with:<ol><li>Neutral data inventory.</li><li>Neutral results report.</li><li>Original EPC certificate.</li></ol></li><li>Energy Consumption Data, with:<ol><li>Files, where available, with energy consumption data for the building, which can be used for validation of models and EPC results.</li></ol></li><li>Drawings<ol><li>Building drawings which can be used as an aid for generating the EPC, or for creating dynamic energy consumption&nbsp; models.</li></ol></li><li>Other Data<ol><li>Any other data that can be useful for the purposes of creating or validating an EPC or an energy consumption dynamic model for the building.</li></ol></li><li>Dynamic Model<ol><li>Data to run a dynamic model of the building, if available.</li></ol></li></ol><p>The files have been redacted to exclude confidential information.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

PL09 - Single family house - Zielonka (Poland)

<p>Data files for building: PL09 - Single family house - Zielonka (Poland)</p><p>Languages: Polish, English</p><p>These files are part of the public benchmark repository created as a part of the crossCert EU project.&nbsp;</p><p>This repository contains curated building data, certificate results and, where available, measured performance results. The repository is publicly available so that it can be used as a testbench for new Energy Performance Certificate (EPC) procedures.</p><p>The files are organised in the following folders&nbsp; (note that not all files are always provided):</p><ol><li>Main data&nbsp; and Results, with:<ol><li>Neutral data inventory.</li><li>Neutral results report.</li><li>Original EPC certificate.</li></ol></li><li>Energy Consumption Data, with:<ol><li>Files, where available, with energy consumption data for the building, which can be used for validation of models and EPC results.</li></ol></li><li>Drawings<ol><li>Building drawings which can be used as an aid for generating the EPC, or for creating dynamic energy consumption&nbsp; models.</li></ol></li><li>Other Data<ol><li>Any other data that can be useful for the purposes of creating or validating an EPC or an energy consumption dynamic model for the building.</li></ol></li><li>Dynamic Model<ol><li>Data to run a dynamic model of the building, if available.</li></ol></li></ol><p>The files have been redacted to exclude confidential information.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

PL15 - Single family house - Pilawa (Poland)

<p>Data files for building: PL15 - Single family house - Pilawa (Poland)</p><p>Languages: Polish, English</p><p>These files are part of the public benchmark repository created as a part of the crossCert EU project.&nbsp;</p><p>This repository contains curated building data, certificate results and, where available, measured performance results. The repository is publicly available so that it can be used as a testbench for new Energy Performance Certificate (EPC) procedures.</p><p>The files are organised in the following folders&nbsp; (note that not all files are always provided):</p><ol><li>Main data&nbsp; and Results, with:<ol><li>Neutral data inventory.</li><li>Neutral results report.</li><li>Original EPC certificate.</li></ol></li><li>Energy Consumption Data, with:<ol><li>Files, where available, with energy consumption data for the building, which can be used for validation of models and EPC results.</li></ol></li><li>Drawings<ol><li>Building drawings which can be used as an aid for generating the EPC, or for creating dynamic energy consumption&nbsp; models.</li></ol></li><li>Other Data<ol><li>Any other data that can be useful for the purposes of creating or validating an EPC or an energy consumption dynamic model for the building.</li></ol></li><li>Dynamic Model<ol><li>Data to run a dynamic model of the building, if available.</li></ol></li></ol><p>The files have been redacted to exclude confidential information.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

PL07 - Single family house - Ostrów Mazowiecka (Poland)

<p>Data files for building: PL07 - Single family house - Ostrów Mazowiecka (Poland)</p><p>Languages: Polish, English</p><p>These files are part of the public benchmark repository created as a part of the crossCert EU project.&nbsp;</p><p>This repository contains curated building data, certificate results and, where available, measured performance results. The repository is publicly available so that it can be used as a testbench for new Energy Performance Certificate (EPC) procedures.</p><p>The files are organised in the following folders&nbsp; (note that not all files are always provided):</p><ol><li>Main data&nbsp; and Results, with:<ol><li>Neutral data inventory.</li><li>Neutral results report.</li><li>Original EPC certificate.</li></ol></li><li>Energy Consumption Data, with:<ol><li>Files, where available, with energy consumption data for the building, which can be used for validation of models and EPC results.</li></ol></li><li>Drawings<ol><li>Building drawings which can be used as an aid for generating the EPC, or for creating dynamic energy consumption&nbsp; models.</li></ol></li><li>Other Data<ol><li>Any other data that can be useful for the purposes of creating or validating an EPC or an energy consumption dynamic model for the building.</li></ol></li><li>Dynamic Model<ol><li>Data to run a dynamic model of the building, if available.</li></ol></li></ol><p>The files have been redacted to exclude confidential information.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

GR03 - Single family house - Falhro (Greece)

<p>Data files for building: GR03 - Single family house - Falhro (Greece)</p><p>Languages: Greek, English</p><p>These files are part of the public benchmark repository created as a part of the crossCert EU project.&nbsp;</p><p>This repository contains curated building data, certificate results and, where available, measured performance results. The repository is publicly available so that it can be used as a testbench for new Energy Performance Certificate (EPC) procedures.</p><p>The files are organised in the following folders&nbsp; (note that not all files are always provided):</p><ol><li>Main data&nbsp; and Results, with:<ol><li>Neutral data inventory.</li><li>Neutral results report.</li><li>Original EPC certificate.</li></ol></li><li>Energy Consumption Data, with:<ol><li>Files, where available, with energy consumption data for the building, which can be used for validation of models and EPC results.</li></ol></li><li>Drawings<ol><li>Building drawings which can be used as an aid for generating the EPC, or for creating dynamic energy consumption&nbsp; models.</li></ol></li><li>Other Data<ol><li>Any other data that can be useful for the purposes of creating or validating an EPC or an energy consumption dynamic model for the building.</li></ol></li><li>Dynamic Model<ol><li>Data to run a dynamic model of the building, if available.</li></ol></li></ol><p>The files have been redacted to exclude confidential information.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

BG04 - Single family house - Sofia (Bulgaria)

<p>Data files for building: &nbsp;BG04 - Single family house - Sofia (Bulgaria)</p><p>Languages: Bulgarian, English</p><p>These files are part of the public benchmark repository created as a part of the crossCert EU project.&nbsp;</p><p>This repository contains curated building data, certificate results and, where available, measured performance results. The repository is publicly available so that it can be used as a testbench for new Energy Performance Certificate (EPC) procedures.</p><p>The files are organised in the following folders&nbsp; (note that not all files are always provided):</p><ol><li>Main data&nbsp; and Results, with:<ol><li>Neutral data inventory.</li><li>Neutral results report.</li><li>Original EPC certificate.</li></ol></li><li>Energy Consumption Data, with:<ol><li>Files, where available, with energy consumption data for the building, which can be used for validation of models and EPC results.</li></ol></li><li>Drawings<ol><li>Building drawings which can be used as an aid for generating the EPC, or for creating dynamic energy consumption&nbsp; models.</li></ol></li><li>Other Data<ol><li>Any other data that can be useful for the purposes of creating or validating an EPC or an energy consumption dynamic model for the building.</li></ol></li><li>Dynamic Model<ol><li>Data to run a dynamic model of the building, if available.</li></ol></li></ol><p>The files have been redacted to exclude confidential information.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

AT02 - Single family house - Schönkirchen (Austria)

<p>Data files for building: &nbsp;AT02 - Single family house - Schönkirchen (Austria)</p><p>Languages: German, English</p><p>These files are part of the public benchmark repository created as a part of the crossCert EU project.&nbsp;</p><p>This repository contains curated building data, certificate results and, where available, measured performance results. The repository is publicly available so that it can be used as a testbench for new Energy Performance Certificate (EPC) procedures.</p><p>The files are organised in the following folders&nbsp; (note that not all files are always provided):</p><ol><li>Main data&nbsp; and Results, with:<ol><li>Neutral data inventory.</li><li>Neutral results report.</li><li>Original EPC certificate.</li></ol></li><li>Energy Consumption Data, with:<ol><li>Files, where available, with energy consumption data for the building, which can be used for validation of models and EPC results.</li></ol></li><li>Drawings<ol><li>Building drawings which can be used as an aid for generating the EPC, or for creating dynamic energy consumption&nbsp; models.</li></ol></li><li>Other Data<ol><li>Any other data that can be useful for the purposes of creating or validating an EPC or an energy consumption dynamic model for the building.</li></ol></li><li>Dynamic Model<ol><li>Data to run a dynamic model of the building, if available.</li></ol></li></ol><p>The files have been redacted to exclude confidential information.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
edi44/100

Data for 'Floral color and family drive contrasting plant-pollinator responses to nutrient enrichment' by Rebecca A. Nelson, Elizabeth T. Borer, and Eric W. Seabloom 2025. Collected in California grasslands 2023 and 2024.

Data for analysis of how flower color and family mediate plant-pollinator response to nutrient enrichment. Data on pollinator visitation and flower abundance were collected in three California grasslands in 2023 and 2024 from a factorial experimental in which combinations of nitrogen, phosphorus, and potassium with micronutrients were applied.

openCustomMay 2025View details →
zenodo40/100

Gene family data from the PhyloGenes (release version 1.2, phylogenes.org)

<p>The compressed file contains:&nbsp;</p> <p><br> 1. PhyloXML_files&nbsp;</p> <p>This folder has family trees in PhyloXML format, one file per family (e.g. &lt;family_ID&gt;.xml).</p> <p>The following information is provided for each node of a tree:<br> 1) leaf node:<br> branch length<br> name &lt;gene_id&gt;<br> taxonomy scientific_name<br> sequence accession &lt;UniProt ID&gt;</p> <p>2) non-leaf&nbsp;node:<br> branch length<br> events &lt;duplication or speciation&gt;</p> <p><br> 2. phylogenes_csv.tar.xz</p> <p>This tar file has gene information of family members in CSV format, one file per family (e.g. &lt;family_ID&gt;.csv).&nbsp;</p> <p>A CSV file includes the following columns:<br> Uniprot ID<br> Gene &lt;Gene name. If none then Gene ID&gt;<br> Gene ID<br> Gene name<br> Organism<br> Subfamily name</p> <p>Any columns displayed after &#39;Subfamily name&#39; are &#39;Known functions&#39;. Each &#39;Known function&#39; is a GO molecular function term that is annotated to at least one member of the gene family AND that the annotation is supported by an experimental evidence. Number 1 or 0 indicates the presence or absence of a particular function in a gene.</p>

opencc-by-4.0Dec 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record