Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,283
datasets available to search
ShareScore release 0.7.1
Dataset results
4,283 results for “Database”
CLDF dataset derived from Mamta's "South Asian Numerals Database" from 2024
<p>Cite the source of the dataset as:</p> <blockquote> <p>Mamta, K. (2024): South Asian Numerals Database (SAND). Leipzig: Max Planck Institute for Evolutionary Anthropology.</p> </blockquote>
Global Transmission Database
<p>The Global Transmission Database (GTD) consists of comprehensive data regarding existing and planned cross-border transmission capacities globally collated from public sources. The dataset is oriented towards representing entry level capacity data (MW) that can be used in energy system models and other computational tools. Transmission capacities are provided at a country-to-country basis in addition to regional level data for a sample of larger countries (Australia, Brazil, Canada, China, India, Indonesia, Japan, Philippines, Russian Federation, United States, Vietnam). Capacities are provided for land-based transmission pathways as well as for subsea pathways. Refer to the accompanying paper for details on the applied methodology as well as a file by file description of the repository content.</p>
Board Leadership Database (U.S. Public Firms) + ML Script for Scaling Human Coded Data
<p>Files include: (1) an open sourced database of CEO duality and board chair orientations developed by scaling human coded data using supervised machine learning techniques (in both .dta and .csv formats), as well as (2) the accompanying training and scoring scripts to scale human coded data.</p> <p>Users may apply the scoring script to score the same variables from company proxy statements, or may adapt the training/scoring scripts and retrain models to scale human coded data of other constructs or measures. </p> <p>We note that early steps in the process to develop our database and script required web-scraping of company filings from SEC Edgar and text extraction from collected filings. We relied on other publicly available scripts to develop our own fetcher and extraction scripts. Users seeking to duplicate those parts of the process may benefit from the following resources from Kai Chen and pipy.org: </p> <p>For resources from Kai Chen: see <a href="https://urldefense.com/v3/__https:/www.kaichen.work/?p=681__;!!K6Z8K8YTIA!BizP9-ZnzgV0Pq7ck-UENJ1EBDrFkAkNoCaO34Ad1ezxH_okstsfxniagvxXdWudkDP44fPACbY9vaIG5A$">https://www.kaichen.work/?p=681</a> and <a href="https://urldefense.com/v3/__https:/www.kaichen.work/?p=946__;!!K6Z8K8YTIA!BizP9-ZnzgV0Pq7ck-UENJ1EBDrFkAkNoCaO34Ad1ezxH_okstsfxniagvxXdWudkDP44fPACbZhp5N_Tw$">https://www.kaichen.work/?p=946</a></p> <p>For resources from pipy.org, see <a href="https://urldefense.com/v3/__https:/pypi.org/project/sec-edgar-downloader/__;!!K6Z8K8YTIA!BizP9-ZnzgV0Pq7ck-UENJ1EBDrFkAkNoCaO34Ad1ezxH_okstsfxniagvxXdWudkDP44fPACbYxYSZGGw$">sec-edgar-downloader</a> and <a href="https://urldefense.com/v3/__https:/pypi.org/project/sec-api/__;!!K6Z8K8YTIA!BizP9-ZnzgV0Pq7ck-UENJ1EBDrFkAkNoCaO34Ad1ezxH_okstsfxniagvxXdWudkDP44fPACbZ_FwY6jw$">sec-api</a></p> <p> </p>
glottolog/glottolog: Glottolog database 5.2.1 as CLDF
<p>Cite the source of the dataset as:</p> <blockquote> <p>Hammarström, Harald & Forkel, Robert & Haspelmath, Martin & Bank, Sebastian. 2025. Glottolog 5.2.1. Leipzig: Max Planck Institute for Evolutionary Anthropology. (Available online at https://glottolog.org)</p> </blockquote>
glottolog/glottolog: Glottolog database 5.2.1
<p>Hammarström, Harald & Forkel, Robert & Haspelmath, Martin & Bank, Sebastian. 2025. Glottolog 5.2.1. Leipzig: Max Planck Institute for Evolutionary Anthropology. (Available online at <a href="https://glottolog.org">https://glottolog.org</a>)</p>
INDIGO Graffiti Reference Database
<p>There is an extensive body of popular and scholarly literature on ancient and recent graffiti. The <strong>INDIGO Graffiti Reference Database</strong> was created to collect much of this graffiti-related literature in one place. It also contains literature about street art because contemporary graffiti and street art are hard to disentangle, and their relationship and hierarchy depend upon the source one consults. The database is a product of the academic graffiti project INDIGO.</p> <p>This collection of references is made available in two forms:</p> <ul> <li>a closed-source <a href="https://www.citavi.com/en">Citavi</a> database (the *.ctv6archive file);</li> <li>an open-source <a href="https://www.zotero.org">Zotero</a> database <ul> <li>either <a href="https://www.zotero.org/groups/5192206/indigo_graffiti_reference_database/library">consultable online</a> (reading only),</li> <li>or downloadable here as a Zotero RDF or BibTex file (*.rdf and *.bib, respectively) for import into Zotero or other reference managers.</li> </ul> </li> </ul>
Pulotu: Database of Austronesian Religions
<p>Cite the source of the dataset as:</p> <blockquote> <p>Watts J., Sheehan O., Greenhill S.J., Gomes-Ng S., Atkinson Q.D., Bulbulia J., Gray R.D. (2015). Pulotu: Database of Austronesian Supernatural Beliefs and Practices. PLoS ONE 10(9). DOI: https://doi.org/10.1371/journal.pone.0136783</p> </blockquote>
SeMRA Gene Mappings Database
<p>Analyze the landscape of gene nomenclature resources, species-agnostic. See instructions for reproduction and usage in the attached README.md.</p>
SeMRA Protein Complex Mappings Database
<p>Analyze the landscape of protein complex nomenclature resources, species-agnostic. See instructions for reproduction and usage in the attached README.md.</p>
SeMRA Anatomy Mappings Database
<p>Supports the analysis of the landscape of anatomy nomenclature resources. See instructions for reproduction and usage in the attached README.md.</p>
SeMRA Cell and Cell Line Mappings Database
<p>Originally a reproduction of the EFO/Cellosaurus/DepMap/CCLE scenario posed in the Biomappings paper, this configuration imports several different cell and cell line resources and identifies mappings between them. See instructions for reproduction and usage in the attached README.md.</p>
SeMRA Disease Mappings Database
<p>Supports the analysis of the landscape of disease nomenclature resources. See instructions for reproduction and usage in the attached README.md.</p>
Let the giants be named - Taxonomic description database for the isopod genus Bathynomus for the DELTA system
<p><span>Morphological characters have been extracted from recent relevant literature on <em>Bathynomus</em>, mainly from the six previously known species from the Atlantic Ocean and gulf of Mexico </span><span>(Magalhães and Young 2003; Lowry and Dempsey 2006; Shipley et al. 2016; Huang et al. 2022)</span><span>. To simultaneously maximize compactness and taxonomic value of the species description, mainly characters that have previously been interpreted as diagnostic at the species level have been included, omitting others which played no role in species delimitation. A taxonomic database has been newly set up using DELTA </span><span>(Dallwitz 1980, 1993; Dallwitz et al. 2006)</span><span>. Character states for all presently known Atlantic congeners have been scored to generate a natural language description and a new identification key </span><span>(Dallwitz 1974)</span><span> for the region as well as a species diagnosis. The species diagnosis has been composed based on the characters used in the identification key.</span></p> <p> </p>
Update of the Xylella spp. host plant database
<p>Following a request from the European Commission, in 2018 EFSA released a renovated database of host plant species of <em>Xylella</em> spp. (<em>including both species</em> <em>X. fastidiosa </em>and <em>X. taiwanensis</em><em>) together with a scientific report</em> (EFSA, 2018). EFSA was tasked to maintain and update this database periodically. The mandate now covers the period 2021-2026 and EFSA is requested to release an update of the database twice per year.</p> <p>In July 2025 EFSA released the twelfth update of the <em>Xylella</em> spp. host plant database (VERSION 12) with information retrieved from literature search up to December 2024 and recent Europhyt outbreak notifications (EFSA, 2025). The protocol applied for the extensive literature review, data collection and reporting, as well as results and lists of host plants are described in detail in the related scientific report (EFSA, 2025).</p> <p>The overall number of <em>Xylella</em> spp. host plants determined with at least two different detection methods or positive with one method (between: sequencing, pure culture isolation) reaches now 463 plant species, 210 genera and 71 families (category A – see section 2.4.2 of EFSA (2025)). Such numbers rise to 727 plant species, 319 genera and 91 families if considered regardless of the detection method applied (category E, see section 2.4.2 of EFSA (2025)).</p> <p>The Excel files here attached represent the VERSION 12 of the <em>Xylella</em> spp. host plants database. For a detailed description of the information included in the database, please consult the related scientific report (EFSA, 2025).</p> <p>The Excel file “<em>Xylella</em> spp. host plants database – VERSION 12” contains several sheets: the LEGENDA (with extensive description of each table), the full detailed raw data of the <em>Xylella</em> spp. host plant database (sheet “observation”) and several examples of data extraction.</p> <p>Additional Excel files contain the lists of host plant species of <em>X. fastidiosa</em> (subsp. unknown (i.e. not reported), <em>fastidiosa</em>, <em>multiplex</em>, <em>pauca</em>, <em>morus</em>, <em>sandyi</em>, <em>tashke</em>, <em>fastidiosa/sandyi</em>) and <em>X. taiwanensis</em> infected naturally, artificially and in not specified conditions, and according to different categories (A, B, C, D, E – see section 2.4.2 of EFSA (2025)). The Excel file “new_host_plant_species_v12” contain the list of new host plant species added to the database in this new update.</p> <p><strong>Question number: EFSA-Q-2025-00045</strong></p> <p><strong>Output number: EN-9564</strong></p> <p><strong>Contacts: plants@efsa.europa.eu</strong></p> <p><em>Bibliography:</em></p> <p>EFSA (European Food Safety Authority). (2018). Scientific report on the update of the <em>Xylella</em> spp. host plant database. <em>EFSA Journal 2018</em>, <em>16</em>(9), 5408, 87 pp. <a href="https://doi.org/10.2903/j.efsa.2018.5408">https://doi.org/10.2903/j.efsa.2018.5408</a> </p> <p>EFSA (European Food Safety Authority), Cavalieri, V., Fasanelli, E., Furnari, G., Gibin, D., Gutierrez Linares, A., La Notte, P., Pasinato, L., & Stancanelli, G. (2025). Update of the <em>Xylella</em> spp. host plant database – Systematic literature search up to 31 December 2024. <em>EFSA Journal</em>, <em>23</em>(7), e9563. <a href="https://doi.org/10.2903/j.efsa.2025.9563">https://doi.org/10.2903/j.efsa.2025.9563</a></p>
CLDF dataset derived from Wichmann et al.'s "ASJP Database" v21 from 2025
<p>Cite the source of the dataset as:</p> <blockquote> <p>Wichmann, Søren, Eric W. Holman, Cecil H. Brown, Matthew S. Dryer, and Qibin Ran (eds.). 2025. The ASJP Database (version 21).</p> </blockquote>
EACTS Adult Cardiac Database (ACD)
<div> <h2>EACTS is the home for the global cardiothoracic surgical community.</h2> </div> <p>We exist to improve outcomes for patients with heart and lung conditions by supporting the global surgical cardiothoracic community and informing best practice with first-class education, cutting-edge learning opportunities, world-renowned journals and publications and pioneering research.</p> <p><a href="https://www.eacts.org/" rel="nofollow">https://www.eacts.org/</a></p> <div> <h3>About the ACD</h3> </div> <p>The EACTS Adult Cardiac Database (ACD) is a collaborative registry and benchmarking tool of cardiac surgical data for centres on an international scale, giving surgical teams the advanced data tools and insights to continually improve outcomes for patients.</p> <div> <h3>Register your interest</h3> </div> <p><a href="https://airtable.com/appRF0GOfjFh6RtRS/pagHKi67AmcK76r2R/form">https://airtable.com/appRF0GOfjFh6RtRS/pagHKi67AmcK76r2R/form</a></p> <p> </p>
PHI-base: the Pathogen-Host Interactions Database, version 5.1
<p><strong>Download the dataset here: <a title="Download PHI-base 5.1" href="https://zenodo.org/records/16738930/files/phi-base_v5.1.zip?download=1">phi-base_v5.1.zip</a></strong></p> <p>The Pathogen–Host Interactions Database (PHI-base) is an online database that catalogues experimentally-verified pathogenicity, virulence and effector genes from fungal, oomycete, and bacterial pathogens, which infect animal, plant, fungal, and insect hosts. PHI-base is a valuable resource in the discovery of genes in medically and agronomically important pathogens, which may be potential targets for chemical intervention.</p> <p>Information in PHI-base is manually curated by domain experts and is supported by strong experimental evidence (for example, gene disruption and gene complementation experiments), as well as references to the literature in which the original experiments are described. Annotations are made using terms from ontologies and controlled vocabularies, including the <a href="https://www.geneontology.org/">Gene Ontology</a> (GO), <a href="https://pubmed.ncbi.nlm.nih.gov/21030441/">Brenda Tissue Ontology</a> (BTO), and the <a href="https://obofoundry.org/ontology/phipo.html">Pathogen–Host Interaction Phenotype Ontology</a> (PHIPO).</p> <p>PHI-base 5 includes data that was curated using a new curation process described in <a href="https://doi.org/10.7554/eLife.84658">Cuzick et. al</a> (2023). Data releases for PHI-base 5 do not use the same schema as data releases from PHI-base 4, but all data records from PHI-base 4 that can be made compatible with the new schema are included with this release. Data releases from PHI-base 4 and PHI-base 5 will occur in parallel until such time that all data from PHI-base 4 can be migrated to PHI-base 5. The PHI-base 4 data releases are available on Zenodo at <a href="https://zenodo.org/doi/10.5281/zenodo.5356870">https://zenodo.org/doi/10.5281/zenodo.5356870</a>.</p> <p>For more information about the planned transition from PHI-base 4 to PHI-base 5, see the <a href="https://phi5.phi-base.org/#/help">Help</a> and <a href="https://phi5.phi-base.org/#/announcements">Announcements</a> page on the PHI-base 5 website.</p> <h2>Release statistics</h2> <p>This version of the PHI-base 5 dataset contains the following types of information:</p> <table style="border-collapse: collapse; border-width: 1px; width: 40.2168%; height: 362.8px;"> <thead> <tr style="height: 19.6px;"> <th style="border-width: 1px; width: 84.846%; height: 19.6px;">Data type</th> <th style="border-width: 1px; width: 15.4054%; height: 19.6px;">Count</th> </tr> </thead> <tbody> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Genes</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">9457</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Interactions</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">31094</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Pathogen species</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">303</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Host species</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">237</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Diseases</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">343</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">References</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">5202</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;"><strong>Annotations</strong></td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;"> </td> </tr> <tr style="height: 10px;"> <td style="border-width: 1px; width: 84.846%; height: 10px;">Pathogen-host interaction phenotype</td> <td style="border-width: 1px; width: 15.4054%; height: 10px;">18260</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Gene-for-gene phenotype</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">452</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Pathogen phenotype</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">9413</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Host phenotype</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">14</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">GO biological process</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">1453</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">GO cellular component</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">85</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">GO molecular function</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">152</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Post-translational modification</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">6</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Physical interaction</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">53</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">WT RNA expression</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">36</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">WT protein expression</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">2</td> </tr> </tbody> </table> <h2>File contents</h2> <ul> <li> <p><strong>phi-base_v5.1.xlsx</strong>: the PHI-base dataset as an Excel spreadsheet. This format follows the layout of the PHI-base 5 website, with sheets corresponding to the sections of gene pages on the website. This format is designed for use by non-technical users.</p> </li> <li> <p><strong>phi-base_v5.1.json</strong>: the PHI-base dataset in JSON format. This is modelled on the export format used by PHI-Canto, the curation tool used by PHI-base. This format is primarily intended for programmatic usage and has additional information (e.g. metadata for curation sessions) that is not included in the spreadsheet format.</p> </li> <li> <p><strong>phi-base.schema.json</strong>: a <a href="https://json-schema.org/">JSON Schema</a> file for the JSON format of the dataset. This is included as documentation for the fields in the JSON file, but can also be used to validate the dataset.</p> </li> </ul>
Updated database of craters on Mars with pitted impact deposits
<p>This point-based database (provided in two formats: as a ESRI shapefile set and simple .csv) currently contains 309 craters on Mars that possess “crater-related pitted materials” (CRPM), which are consistent with impact deposits described in detail in Tornabene et al. (2007; 2012). The Tornabene et al. (2012) publication is the original source of the initial database of 204 craters, which was based on a survey by the Mars Reconnaissance Orbiter (MRO) over a period of late 2006 to early 2012. MRO continues to image the surface and its craters, as such the database has since grown from 204 to 309 entries to date. Despite this growth, the general characteristics of the crater population remains generally consistent with what is described in Tornabene et al. (2012) (e.g., size range, latitudinal and elevation distribution, etc.).</p> <p>When present, these pitted impact deposits represent the upper most surface of the crater-fill with the pits potentially representing top-down views of so-called degassing pipes observed only in eroded cross-sections at some terrestrial impact structures such as the Ries in Germany (e.g., Caudill et al., 2021). Therefore, the craters that contain pits and preserve them well are themselves amongst the very best-preserved and often youngest craters of their size-class on Mars. Indeed, some of these craters are observed to have far-reaching (10s to 100s of crater radii) thermal / secondary crater rays (e.g., Tornabene et al. 2006), which is considered to be a feature associated with only the best-preserved and youthful craters on planets/moons with solid surfaces.</p> <p>These craters have enabled us to place further constraints on the scaling of crater depth as a function of diameter for complex craters on Mars (Tornabene et al. 2018) and may even help us to ultimately determine where the only samples we have of Mars — the Martain Meteorites — come from.</p> <p>See README rtf file for further details on the database.</p> <p> </p> <p><strong>Versions</strong></p> <p><strong>8.21.2025: </strong>4th version - deleted 3 additional duplicates (Lunae, Oudemans and Toro) total entries is now 309<strong><br></strong></p> <p><strong>8.20.2025b</strong>: 3rd version upload - fixed 1 duplicate (312 entries), caught some additional updates with respect to new official crater names, and CTX image IDs.</p> <p><strong>8.20.2025</strong>: 2nd version with an increase to 313 entries with some updates to preservation ratings, image IDs, etc.</p> <p><strong>5.3.2023</strong>: 1st version uploaded with 300 entries</p> <p> </p> <p><strong>Main references (*original/source database):</strong></p> <p>*Tornabene, L.L., Osinski, G.R., McEwen, A.S., Boyce, J.M., Bray, V.J., Caudill, C.M., Grant, J.A., Hamilton, C.W., Mattson, S. and Mouginis-Mark, P.J., 2012. Widespread crater-related pitted materials on Mars: Further evidence for the role of target volatiles during the impact process. Icarus, 220(2), pp.348-368. https://doi.org/10.1016/j.icarus.2012.05.022</p> <p>Tornabene, L.L., McEwen, A.S., Osinski, G.R., Mouginis-Mark, P.J., Boyce, J.M., Williams, R.M.E., Wray, J.J. and Grant, J.A., 2007. Impact melting and the role of subsurface volatiles: Implications for the formation of valley networks and phyllosilicate-rich lithologies on early Mars. In International Conf. on Mars VII. Lunar Planet. Sci. Inst. Contri (Vol. 1353), Abstract# 3288.</p> <p><strong>Other references:</strong></p> <p>Tornabene, L.L., Moersch, J.E., McSween Jr, H.Y., McEwen, A.S., Piatek, J.L., Milam, K.A. and Christensen, P.R., 2006. Identification of large (2–10 km) rayed craters on Mars in THEMIS thermal infrared images: Implications for possible Martian meteorite source regions. Journal of Geophysical Research: Planets, 111(E10).</p> <p>Boyce, J.M., Wilson, L., Mouginis-Mark, P.J., Hamilton, C.W. and Tornabene, L.L., 2012. Origin of small pits in martian impact craters. Icarus, 221(1), pp.262-275.</p> <p>Denevi, B.W., Blewett, D.T., Buczkowski, D.L., Capaccioni, F., Capria, M.T., De Sanctis, M.C., Garry, W.B., Gaskell, R.W., Le Corre, L., Li, J.Y. and Marchi, S., 2012. Pitted terrain on Vesta and implications for the presence of volatiles. Science, 338(6104), pp.246-249.</p> <p>Sizemore, H.G., Platz, T., Schorghofer, N., Prettyman, T.H., De Sanctis, M.C., Crown, D.A., Schmedemann, N., Neesemann, A., Kneissl, T., Marchi, S. and Schenk, P.M., 2017. Pitted terrains on (1) Ceres and implications for shallow subsurface volatile distribution. Geophysical Research Letters, 44(13), pp.6570-6578.</p> <p>Tornabene, L.L., Watters, W.A., Osinski, G.R., Boyce, J.M., Harrison, T.N., Ling, V. and McEwen, A.S., 2018. A depth versus diameter scaling relationship for the best-preserved melt-bearing complex craters on Mars. Icarus, 299, pp.68-83.</p> <p>Caudill, C., Osinski, G.R., Greenberger, R.N., Tornabene, L.L., Longstaffe, F.J., Flemming, R.L. and Ehlmann, B.L., 2021. Origin of the degassing pipes at the Ries impact structure and implications for impact‐induced alteration on Mars and other planetary bodies. Meteoritics & Planetary Science, 56(2), pp.404-422.</p> <p>Michalik, T., Matz, K.D., Schröder, S.E., Jaumann, R., Stephan, K., Krohn, K., Preusker, F., Raymond, C.A., Russell, C.T. and Otto, K.A., 2021. The unique spectral and geomorphological characteristics of pitted impact deposits associated with Marcia crater on Vesta. Icarus, 369, p.114633.</p>
TetrapodTraits Database
<h2>Abstract </h2> <p>Tetrapods (amphibian, reptiles, birds and mammals) are model systems for global biodiversity science, but continuing data gaps, limited data standardisation, and ongoing flux in taxonomic nomenclature constrain integrative research on this group and potentially cause biassed inference. We combined and harmonised taxonomic, spatial, phylogenetic, and attribute data with phylogeny-based multiple imputation to provide a comprehensive data resource (TetrapodTraits 1.0.0) that includes values, predictions, and sources for body size, activity time, micro- and macrohabitat, ecosystem, threat status, biogeography, insularity, environmental preferences and human influence, for all 33,281 tetrapod species covered in recent fully sampled phylogenies. The TetrapodTraits 2.0.0 includes information on diet, longevity, litter/clutch size, developmental mode, biome prevalence, population trend, and major threat types, besides expanding the original data with complementary fields reporting revised and imputed threat status and human influence metrics. While there is an obvious need for further data collection and updates, our phylogeny-informed database of tetrapod traits can support a more comprehensive representation of tetrapod species and their attributes in ecology, evolution, and conservation research.</p> <p><strong>Version 2.0.0</strong> (25 September 2025). In addition to the 24 traits reported in TetrapodTraits v. 1.0.1 (see below), we now include 9 new attributes in the TetrapodTraits v. 2.0.0, as described in our new preprint (<a href="https://doi.org/10.21203/rs.3.rs-7556378/v1" target="_blank" rel="noopener">Pyron et al., 2025</a>) as follows:</p> <ol> <li>Diet – includes 12 attributes covering 10 major diet types (invertebrates, endotherm vertebrates, ectotherm vertebrates, unknown vertebrates, fishes, scavenger, fruits, nectar, seeds, other plant materials) as dummy (binary) variables, data sources, and the sum of diet categories. Diet categories follow those previously reported in the EltonTraits database (<a href="https://doi.org/10.1890/13-1917.1" target="_blank" rel="noopener">Wilman et al., 2014</a>). To ensure consistency across tetrapod classes, we adapted EltonTraits data for birds and mammals by applying a threshold of 20% of a species’ relative usage to indicate presence or absence within each diet category. Diet data are available for 23,594 species (70.9% of data coverage, encompassing at least one representative in 99.8% of families and 96.4% of genera). Field names: DietInv, DietVend, DietVect, DietVfish, DietVunk, DietScav, DietFruit, DietNect, DietSeed, DietPlant, SourceDiet, DietBreadth.</li> <li>Longevity – includes 2 attributes that report maximum longevity (in years) and data sources. Longevity data are available for 7,619 species (22.9% of data coverage, representing at least one species in 85.9% of families and 49.2% of genera). Field names: MaxLongevity, SourceLongevity. </li> <li>Litter Size – includes 3 attributes that describe the average litter or clutch size, the detailed metric reported (average, geometric mean, maximum, midpoint, unique value, or unspecified central tendency metric). Litter/clutch size was typically represented by the midpoint between reported maximum and minimum values or the geometric mean across available samples, although other measures of central tendency or single reported values were also included to maximize data coverage. Litter/clutch size data are available for 18,638 species (56% of data coverage, including at least one representative of 96.3% of families and 86.3% of genera). Field names: LitterSize, LitterSize_Metric, SourceLitterSize.</li> <li>Developmental mode – includes 6 attributes that inform direct development, larval stage, viviparity (all three as binary), and data their respective data sources. Given the deep homology of these traits across major tetrapod lineages, we used taxonomic imputation for missing data, leveraging phylogenetic conservatism (e.g., most amphibians have larval stages, while birds are oviparous). Data on developmental mode are available for 31,583 species (94.9% of data coverage, encompassing all families and 98.8% of genera). Field names: DirectDev, LarvalStage, Viviparity, SourceDirectDev, SourceLarvalStage, SourceViviparity.</li> <li>Biome coverage – includes 14 attributes that reflect the proportion of the species’ range maps overlapped by each WWF biome. This information is available for 33,271 species. Field names: BiomeNum_1 to BiomeNum_14. Data derived from <a href="https://doi.org/10.1093/biosci/bix014" target="_blank" rel="noopener">Dinerstein et al. (2017)</a>.</li> <li>PropGrazingArea – Proportion of species range map covered by Grazing land at year 2017 (equals to the sum of PropPastureArea and PropRangeLandArea). Data derived from <a href="https://www.pbl.nl/en/image/links/hyde">HYDE v. 3.2</a>.</li> <li>PropAnthroHyde – Proportion of species range map covered by human-modified areas at year 2017 (equals to the sum of PropPastureArea, PropCropland, PropUrbanArea). Data derived from <a href="https://www.pbl.nl/en/image/links/hyde">HYDE v. 3.2</a>.</li> <li>Population trend – includes 1 ordinal variable indicating the IUCN population trend (0 = decreasing, 1 = stable, 2 = increasing) as reported by the IUCN v. 2023-1. This information is available for 22,741 species (68.3% of data coverage). Field name: PopulationTrend.</li> <li>Threat type – includes 13 attributes that reflect 12 major threat types corresponding to the first level of the <a href="https://www.iucnredlist.org/resources/threat-classification-scheme" target="_blank" rel="noopener">IUCN Threats Classification Scheme v. 3.3</a>, and an additional variable capturing the total number of threat types recorded for each species. Threat type data are available for 16,876 species. Field names: Threat_1 to Threat_12, ThreatSum.</li> </ol> <p>In a complementary manner, we expanded the number of fields reported for the following traits of TetrapodTraits:</p> <ul> <li>RangeSize_km2 – the field RangeSize was originally reported in the TetrapodTraits v. 1.0.0 rep as the number of 110×110 grid cells covered by the species range map (data derived from MOL database). We now include the field RangeSize_km2 to inform the total area (in km²) covered by each species range mapped in an equal area projection.</li> <li>Threat status – the TetrapodTraits v. 1.0.0 included three fields related to threat status, namely IUCN_Binomial, AssessedStatus, SourceStatus. Those fields included non-Data Deficient (DD) assessed statuses for 29,837 species, and 3,444 species listed as Data Deficient (DD; 2,936 species) or unassessed (UA; 508 species). We have now added five new fields: AssessedStatus2, SourceStatus2, ThreatStatus, ImputedThreatStatus, and SourceImputedStatus. The field AssessedStatus2 mirrors AssessedStatus but replaces the statuses of 130 DD/UA species (according to IUCN v. 2024.3) with prior confident assessments from the IUCN or taxonomic working groups, with SourceStatus2 documenting the respective sources. The field ThreatStatus further builds on AssessedStatus2 by replacing the statuses of 2,221 DD/UA species with available imputed statuses for amphibians (739 species; <a href="https://doi.org/10.1016/j.cub.2019.04.005" target="_blank" rel="noopener">Gonzalez-del-Pliego et al., 2019</a>), squamates (1,387 species; <a href="https://doi.org/10.1371/journal.pbio.3001544" target="_blank" rel="noopener">Caetano et al., 2022</a>), chelonians and crocodilians (60 species; <a href="https://doi.org/10.1186/s12862-020-01642-3" target="_blank" rel="noopener">Colston et al., 2020</a>), and birds (35 species; <a href="https://doi.org/10.1016/j.cub.2014.03.011">Jetz et al., 2014</a>). This expansion increases ThreatStatus data coverage to 32,058 species (96.3%), leaving only 1,093 as DD or UA. The ImputedThreatStatus field indicates whether the ThreatStatus is imputed (1) or not (0), and SourceImputedStatus specifies the source of each imputed status.</li> </ul> <p>The TetrapodTraits 2.0.0 also addresses minor issues identified on previously published attribute data, including the following changes: (i) Updated body mass value for the anuran <em>Adelophryne meridionalis.</em> (ii) The replacement of the imputed body mass for the toad <em>Rhinella achalensis</em> with an observed value and respective data source. (iii) Change in the assessed status of the turtle <em>Pelusios castaneus</em> to follow the assessment reported in <a href="https://iucn-tftsg.org/wp-content/uploads/crm.8.checklist.atlas_.v9.2021.e3.pdf" target="_blank" rel="noopener">Rhodin et al. (2021. Turtles of the World 9th Ed. Chelonian Research Monographs, 8:1–472)</a>. (iv) Update to the geographic information (spatial intersections) reported for the rodent <em>Neotoma bryanti</em> and the bird <em>Pipile pipile</em> to follow the geographic distribution reported in IUCN v. 2023-1, as well as the updated within-range attributes for these two species (RangeSize, Latitude, Longitude, Biogeography, Insularity, AnnuMeanTemp, AnnuPrecip, TempSeasonality, PrecipSeasonality, Elevation, ETA50K, HumanDensity, PropUrbanArea, PropCroplandArea, PropPastureArea, and PropRangelandArea).</p> <p><strong>Version 1.0.1 </strong>(25 May 2024). This minor release addresses a spelling error in the file <em>Tetrapod_360.csv</em>. The error involves replacing white-space characters with underscore characters in the field <em>Scientific.Name</em> to match the spelling used in the file<strong> </strong><em>TetrapodTraits_1.0.0.csv</em>. These corrections affect only 102 species considered extinct and 13 domestic species (Bos_frontalis, Bos_grunniens, Bos_indicus, Bos_taurus, Camelus_bactrianus, Camelus_dromedarius, Capra_hircus, Cavia_porcellus, Equus_caballus, Felis_catus, Lama_glama, Ovis_aries, Vicugna_pacos). All extinct and domestic species in TetrapodTraits have their binomial names separated by underscore symbols instead of white space. Additionally, we have added the file <em>GridCellShapefile.zip</em>, which contains the shapefile required to map species presence across the 110 × 110 km equal area grid cells (this file was previously provided through an External Source <a href="../records/10582070/files/Shapefiles.zip?download=1">here</a>).</p> <p><strong>Version 1.0.0</strong> (19 April 2024). TetrapodTraits, the full phylogenetically coherent database we developed, is being made publicly available to support a range of research applications in ecology, evolution, and conservation and to help minimise the impacts of biassed data in this model system. The database includes 24 species-level attributes linked to their respective sources across 33,281 tetrapod species. Specific fields clearly label data sources and imputations in the TetrapodTraits, while additional tables record the 10K values per missing entry per species.</p> <ol> <li>Taxonomy – includes 8 attributes that inform scientific names and respective higher-level taxonomic ranks, authority name, and year of species description. Field names: Scientific.Name, Genus, Family, Suborder, Order, Class, Authority, and YearOfDescription.</li> <li>Phylogenetic tree – includes 2 attributes that notify which fully-sampled phylogeny contains the species, along with whether the species placement was imputed or not in the phylogeny. Field names: TreeTaxon, TreeImputed.</li> <li>Body size – includes 7 attributes that inform length, mass, and data sources on species sizes, and details on the imputation of species length or mass. Field names: BodyLength_mm, LengthMeasure, ImputedLength, SourceBodyLength, BodyMass_g, ImputedMass, SourceBodyMass.</li> <li>Activity time – includes 5 attributes that describe period of activity (e.g., diurnal, fossorial) as dummy (binary) variables, data sources, details on the imputation of species activity time, and a nocturnality score. Field names: Diu, Noc, ImputedActTime, SourceActTime, Nocturnality.</li> <li>Microhabitat – includes 8 attributes covering habitat use (e.g., fossorial, terrestrial, aquatic, arboreal, aerial) as dummy (binary) variables, data sources, details on the imputation of microhabitat, and a verticality score. Field names: Fos, Ter, Aqu, Arb, Aer, ImputedHabitat, SourceHabitat, Verticality.</li> <li>Macrohabitat – includes 19 attributes that reflect major habitat types according to the <a href="https://www.iucnredlist.org/resources/habitat-classification-scheme">IUCN classification</a>, the sum of major habitats, data source, and details on the imputation of macrohabitat. Field names: MajorHabitat_1 to MajorHabitat_10, MajorHabitat_12 to MajorHabitat_17, MajorHabitatSum, ImputedMajorHabitat, SourceMajorHabitat. MajorHabitat_11, representing the marine deep ocean floor (unoccupied by any species in our database), is not included here.</li> <li>Ecosystem – includes 6 attributes covering species ecosystem (e.g., terrestrial, freshwater, marine) as dummy (binary) variables, the sum of ecosystem types, data sources, and details on the imputation of ecosystem. Field names: EcoTer, EcoFresh, EcoMar, EcosystemSum, ImputedEcosystem, SourceEcosystem.</li> <li>Threat status – includes 3 attributes that inform the assessed threat statuses according to IUCN red list and related literature. Field names: IUCN_Binomial, AssessedStatus, SourceStatus.</li> <li>RangeSize – the number of 110×110 grid cells covered by the species range map. Data derived from <a href="https://mol.org/">MOL</a>.</li> <li>Latitude – coordinate centroid of the species range map.</li> <li>Longitude – coordinate centroid of the species range map.</li> <li>Biogeography – includes 8 attributes that present the proportion of species range within each WWF biogeographical realm. Field names: Afrotropic, Australasia, IndoMalay, Nearctic, Neotropic, Oceania, Palearctic, Antarctic. Data derived from <a href="https://doi.org/10.1093/biosci/bix014" target="_blank" rel="noopener">Dinerstein et al. (2017)</a>.</li> <li>Insularity – includes 2 attributes that notify if a species is insular endemic (binary, 1 = yes, 0 = no), followed by the respective data source. Field names: Insularity, SourceInsularity.</li> <li>AnnuMeanTemp – Average within-range annual mean temperature (Celsius degree). Data derived from <a href="https://chelsa-climate.org/">CHELSA v. 1.2</a>.</li> <li>AnnuPrecip – Average within-range annual precipitation (mm). Data derived from <a href="https://chelsa-climate.org/">CHELSA v. 1.2</a>.</li> <li>TempSeasonality – Average within-range temperature seasonality (Standard deviation × 100). Data derived from <a href="https://chelsa-climate.org/">CHELSA v. 1.2</a>.</li> <li>PrecipSeasonality – Average within-range precipitation seasonality (Coefficient of Variation). Data derived from <a href="https://chelsa-climate.org/">CHELSA v. 1.2</a>.</li> <li>Elevation – Average within-range elevation (metres). Data derived from topographic layers in <a href="https://www.earthenv.org/topography">EarthEnv</a>.</li> <li>ETA50K – Average within-range estimated time to travel to cities with a population >50K in the year 2015. Data from <a href="https://doi.org/10.1038/s41597-019-0265-5">Nelson et al. (2019)</a>.</li> <li>HumanDensity – Average within-range human population density in 2017. Data derived from <a href="https://www.pbl.nl/en/image/links/hyde">HYDE v. 3.2</a>.</li> <li>PropUrbanArea – Proportion of species range map covered by built-up area, such as towns, cities, etc. at year 2017. Data derived from <a href="https://www.pbl.nl/en/image/links/hyde">HYDE v. 3.2</a>.</li> <li>PropCroplandArea – Proportion of species range map covered by cropland area, identical to FAO's category 'Arable land and permanent crops' at year 2017. Data derived from <a href="https://www.pbl.nl/en/image/links/hyde">HYDE v. 3.2</a>.</li> <li>PropPastureArea – Proportion of species range map covered by cropland, defined as Grazing land with an aridity index > 0.5, assumed to be more intensively managed (converted in climate models) at year 2017. Data derived from <a href="https://www.pbl.nl/en/image/links/hyde">HYDE v. 3.2</a>.</li> <li>PropRangelandArea – Proportion of species range map covered by rangeland, defined as Grazing land with an aridity index < 0.5, assumed to be less or not managed (not converted in climate models) at year 2017. Data derived from <a href="https://www.pbl.nl/en/image/links/hyde">HYDE v. 3.2</a>.</li> </ol> <p><strong>Additional Information:</strong> This work is output of the <a href="https://vertlife.org/data/">VertLife</a> project. To flag erros, provide updates, or leave other comments, please go to <a href="https://vertlife.org/">vertlife.org</a>. We aim to develop the database into a living resource at <a href="https://vertlife.org/">vertlife.org</a> and your feedback is essential to improve data quality and support community use.</p> <h2>File content</h2> <p>All files use UTF-8 encoding.</p> <ul> <li><strong>TetrapodTraits_2.0.0.csv </strong>–<strong> </strong> the complete TetrapodTraits database, with missing data entries in natural history traits (body length, body mass, activity time, and microhabitat) replaced by the average across the 10K imputed values obtained through phylogenetic multiple imputation. Please note that imputed microhabitat (attribute fields: Fos, Ter, Aqu, Arb, Aer) and imputed activity time (attribute fields: Diu, Noc) are continuous variables within the 0-1 range interval. At the user's discretion, the types of microhabitat and activity time can be transformed into binary variables using a predefined threshold (e.g., 0.50), although we recommend utilizing the original imputed values.<br><br></li> <li><strong>Tetrapod_360_2.0.0.csv </strong>–<strong> </strong>spatial intersections of the 110 x 110 km quadrats shapefile (<em>GridCellShapefile.zip</em>) with species geographic range maps from <a href="https://mol.org">https://mol.org</a>, following updates reported in TetrapodTraits 2.0.0.<br><br></li> <li><strong>GridCellShapefile.zip</strong> –<strong> </strong>contains grid cell shapefiles with a spatial resolution of 110 km, which are required to map the species listed in the <em>Tetrapod_360.csv</em> file. Please note that due to the limitation on the number of characters in shapefile field names, the field names of gridcells_110km.shp are displayed as ("Cl_I110", "Long", "Lat", "WWF_Rlm", "PrpLndA"). Be aware to rename field names to ("Cell_Id110", "Long", "Lat", "WWF_Realm", "PropLandArea") to match the terminology used for the 110 x 110 km grid cells in other files.<br><br></li> <li><strong>ImputedSets.zip </strong>–<strong> </strong>the phylogenetic multiple imputation framework applied to the TetrapodTraits database produced 10,000 imputed values per missing data entry (= 100 phylogenetic trees x 10 validation-folds x 10 multiple imputations). These imputations were specifically developed for four fundamental natural history traits: Body length, Body mass, Activity time, and Microhabitat. To facilitate the evaluation of each imputed value in a user-friendly format, we offer 10,000 tables containing both observed and imputed data for the 33,281 species in the TetrapodTraits database. Each table encompasses information about the four targeted natural history traits, along with designated fields (e.g., ImputedMass) that clearly indicate whether the trait value provided (e.g., BodyMass_g) corresponds to observed (e.g., ImputedMass = 0) or imputed (e.g., ImputedMass = 1) data. Given that the complete set of 10,000 tables necessitates nearly 17GB of storage space, we have organized sets of 1,000 tables into separate zip files to streamline the download process.<br> <ul> <li>ImputedSets_1K.zip, imputations for trees 1 to 10.</li> <li>ImputedSets_2K.zip, imputations for trees 11 to 20.</li> <li>ImputedSets_3K.zip, imputations for trees 21 to 30.</li> <li>ImputedSets_4K.zip, imputations for trees 31 to 40.</li> <li>ImputedSets_5K.zip, imputations for trees 41 to 50.</li> <li>ImputedSets_6K.zip, imputations for trees 51 to 60.</li> <li>ImputedSets_7K.zip, imputations for trees 61 to 70.</li> <li>ImputedSets_8K.zip, imputations for trees 71 to 80.</li> <li>ImputedSets_9K.zip, imputations for trees 81 to 90.</li> <li>ImputedSets_10K.zip, imputations for trees 91 to 100.</li> </ul> </li> </ul> <h2>External files</h2> <p>The R-code used for data analysis is available at <a href="../records/10976274">10.5281/zenodo.10582069</a>.</p> <h2>Funding</h2> <p>São Paulo Research Foundation (FAPESP) for grants supporting MRM (#2021/11840-6 and #2022/12231-6), LFT (#2016/25358-3), KC (#2020/12558-0), and RZC (#2022/15247-0); Coordenação de Aperfeiçoamento de Pessoal de Nível Superior (CAPES) for the fellowship to JJMG; Conselho Nacional de Desenvolvimento Científico - CNPq for research grants in support of FPW (#311504/2020-5) and LFT (#302834/2020-6); U.S. National Science Foundation (NSF) for grants supporting RAP (DEB-1441719), RCKB (DEB-1441652), and WJ (DEB-1441737 and DEB-1441719). WJ also acknowledges support from NASA grants 80NSSC17K0282 and 80NSSC18K0435; and the E.O. Wilson Biodiversity Foundation.</p> <h2>Citation</h2> <p>When using data from TetrapodTraits 2.0.0, please cite both the following foundational works:</p> <p>Pyron, R.A., Moura, M.R., Bowie, R.K., Brito, S.F., Ceron, K., Colston, T.J., Esselstybm J.A., Guedes, J.J.M., Guralnick, R.P., Hidalgo, V.G.L., Mooers, A., Moroti, M.T., Paiva, M.F., Pennel, M., Pirani, R.M., Souza, J.A.S., Upham, N.S., Xavier, J., Jetz, W. <strong>Anthropocene Imperilment of Ancient Diversity and Evolutionary Potential inTerrestrial Vertebrates</strong>. <em>Research Square</em> (preprint), 2025. doi: <a href="https://doi.org/10.21203/rs.3.rs-7556378/v1" target="_blank" rel="noopener">10.21203/rs.3.rs-7556378/v1</a></p> <p>Moura, M.R., Ceron, K., Guedes, J.J.M., Chen-Zhao, R., Sica, Y.V., Hart, J., Dorman, W., Gonzalez-del-Pliego, P., Ranipeta, A., Catenazzi, A.., Werneck, F.P., Toledo, L.F., Upham, N.S., Tonini, J.F.R., Colston, T.J., Guralnick, R., Bowie, R.C.K., Pyron, R.A., Jetz, W. <strong>A phylogeny-informed characterisation of global tetrapod traits addresses data gaps and biases</strong>. <em>PLoS Biology</em>, 22:e3002658, 2024. doi: <a href="https://doi.org/10.1371/journal.pbio.3002658">10.1371/journal.pbio.3002658</a></p> <p>Correspondence to: mariormoura@gmail.com</p>
D3.3. Experimental Wave-Tank Validation Database
<p>This database is a deliverable of the FLOATECH project, funded under the European Union's Horizon 2020 research and innovation program under grant agreement No 101007142.</p><p>The aim of the accompanying document is to describe the experimental testing campaign C2 at the LHEEA wave-tank facility. The campaign took place between May and June 2023. The objective of the campaign is to test several FOWT control strategy, including a feed forward wave-based control, using the software-in-the-loop SOFTWIND system. The accompanying report on the E.U. portal describes the database created from these experiments and aimed to be shared for model validation. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.