Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,088

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,088 results for “worldwide”

Learn how ShareScore rates datasets ↗
edi60/100

Modeling Impacts of Climate Change on Mangroves Worldwide 2012-2080

Given the multitude of ecosystem services provided by mangroves, it is important to understand their potential responses to global climate change. Extensive reviews of the literature and manipulative experiments suggest that mangroves will be impacted by climate change, but few studies have tested these predictions over large scales using statistical models. We provide the first example of applying species and community distribution models (SDMs and CDMs, respectively) to coastal mangroves worldwide. Species projected to shift their ranges polewards by at least 2 degrees of latitude consistently experience a decrease in the amount of suitable coastal area available to them. Central America and the Caribbean are forecast to lose more mangrove species than other parts of the world. We found that the extent and grain size, at which continuous CDM outputs are examined, independent of the grain size at which the models operate, can dramatically influence the number of pseudo-absences needed for optimal parameterization. The SDMs and CDMs presented here provide a first approximation of how mangroves will respond to climate change given simple correlative relationships between occurrence records and environmental data. Additional, precise georeferenced data on mangrove localities and concerted efforts to collect data on ecological processes across large-scale climatic gradients will enable future research to improve upon these correlative models.

openCC0Dec 2023View details →
edi60/100

Root Systems of Individual Plants Worldwide 2022

The above- and below-ground sizes and shapes of plants strongly influence plant competition, community structure, and plant-environment interactions, but the plant size and shape across climate regimes remain incompletely understood. In this study I seek to understand how plant geometries respond to varying climates via trade-offs in shoot height and width, and root depth and spread. I more than doubled the Root Systems of Individual Plants (RSIP) database to contain 5,647 observations, to our knowledge the largest database describing the maximum rooting depth, lateral spread, and shoot size of terrestrial plants in the world. Shoot size and root system size strongly covary. Across climatic gradients woody plants show deeper-narrower root systems in arid climates and taller shoots in humid climates. Phylogeny greatly influences shoot size. Rooting depth is primarily influenced by climate seasonality and lateral root spread is strongly influenced by shoot size. Using our newly expanded global database I found that shoot size covaries strongly with rooting system size; however, these relationships are not static across the climate space, as the geometries of plants shift considerably.

openCC0Dec 2023View details →
edi56/100

Prey Capture by Carnivorous Plants Worldwide 1923-2007

Available phylogenetic data illustrate that in all carnivorous lineages, the ancestral trap type is a sticky, flypaper-type trap (Ellison and Gotelli, 2001). In the Caryophyllales, pitfall traps (Nepenthes) and snap traps (Dionaea and Aldrovanda) are derived relative to the sticky pads of Drosera. Similarly, in the Lamiales, the sticky-leaved Pinguicula is ancestral to Genlisea with its eel (or lobster-pot) traps and Utricularia with its vacuum traps. In the Ericales, the Sarraceniaceae with its pitfall traps are derived relative to Roridula, another species with flypaper traps. Muller et al. (2004) hypothesed that carnivorous genera with rapidly evolving genomes (Genlisea and Utricularia) have more predictable and frequent captures of prey than do genera with more slowly evolving genomes; by extension it could be hypothesized that in general, carnivorous plants with more complex traps should have more predictable and frequent captures of prey than do those with relatively simple traps. Increases in predictability and frequency of prey capture could be achieved by evolving more elaborate mechanisms for attracting prey, by specializing on particular types of prey, or, as Darwin suggested, by specializing on particular (large) sizes of prey. In all cases, one would expect that prey actually captured would not be a random sample of the available prey. Furthermore, when multiple species of carnivorous plants co-occur, one would predict, again following Darwin that interspecific competition would lead to specialization on particular kinds of prey. Because the traps of carnivorous plants accumulate identifiable remains of prey, analysis of trap contents can provide an aggregate record of the prey that have been successfully "sampled" by the plant. Such samples could be used to begin to test the hypothesis that carnivorous plant genera differ in prey composition and to look for evidence of specialization in prey capture. Over the past 80 years, numerous ecologists have gat

openCC0Dec 2023View details →
edi56/100

Ecophysiology of Carnivorous Plants Worldwide 1980-2011

Identification of trade-offs among physiological and morphological traits and their use in cost-benefit models and ecological or evolutionary optimization arguments have been hallmarks of ecological analysis for at least 50 years. Carnivorous plants are model systems for studying a wide range of ecophysiological and ecological processes and the application of a cost-benefit model for the evolution of carnivory by plants has provided many novel insights into trait-based cost-benefit models. Central to the cost-benefit model for the evolution of botanical carnivory is the relationship between nutrients and photosynthesis; of primary interest is how carnivorous plants efficiently obtain scarce nutrients that are supplied primarily in organic form as prey, digest and mineralize them so that they can be readily used, and allocate them to immediate versus future needs. Most carnivorous plants are terrestrial - they are rooted in sandy or peaty wetland soils - and most studies of cost-benefit trade-offs in carnivorous plants are based on terrestrial carnivorous plants. However more than 10% of carnivorous plants are unrooted aquatic plants. By examining data published between 1980 and 2011, we ask whether the cost-benefit model applies equally well to aquatic carnivorous plants and what general insights into trade-off models are gained by this comparison. Nutrient limitation is more pronounced in terrestrial carnivorous plants, which also have much lower growth rates and much higher ratio of dark respiration to photosynthetic rates than aquatic carnivorous plants. Phylogenetic constraints on ecophysiological trade-offs among carnivorous plants remain unexplored. Despite differences in detail, the general cost-benefit framework continues to be of great utility in understanding the evolutionary ecology of carnivorous plants. We provide a research agenda that if implemented would further our understanding of ecophysiological trade-offs in carnivorous plants and also would pro

openCC0Dec 2023View details →
edi52/100

GHG-depths: Greenhouse gas depth-profile data in 522 lakes worldwide

Lakes, ponds, and reservoirs (hereafter: “lakes”) are significant sources of the greenhouse gases carbon dioxide (CO2) and methane (CH4). Emissions of CO2 and CH4 from lakes are regulated in part by in-lake processes, including the production and storage of gases in the lower parts of the water column (bottom waters). However, while substantial efforts have been made to improve estimates of greenhouse gas emissions from lakes, limited data on gas concentrations along depth profiles have prevented the incorporation of bottom-water processes in global emission estimates. Here, we present GHG-depths: the largest existing dataset of depth-profile CO2 and CH4 measurements worldwide, including 522 lakes across 38 countries and all seven continents. These data include contributions from 45 research teams and 56 published studies, totaling 2558 discrete sampling events. As global change continues to alter biogeochemical cycling in lakes, these data can help improve mechanistic models to better predict greenhouse gas production and emission from lakes worldwide.

openCC (other)Jan 2026View details →
edi52/100

Soil and root-associated fungal response to nitrogen and phosphorus addition from grasslands worldwide: 2011-2012.

Ecosystems across the globe receive elevated inputs of nutrients, but the consequences of this for soil fungal guilds that mediate key ecosystem functions remain unclear. We found that nitrogen and phosphorus addition to 25 grasslands distributed across four continents promoted the relative abundance of fungal pathogens, suppressed mutualists, but did not affect saprotrophs. Structural equation models suggested that responses were often indirect and primarily mediated by nutrient-induced shifts in plant communities. Nutrient addition also reduced co-occurrences within and among fungal guilds, which could have important consequences for belowground interactions. Focusing only on plots that received no nutrient addition, soil properties influenced pathogen abundance globally, whereas plant community characteristics influenced mutualists, and climate influenced saprotrophs. These guild-level responses enhance our ability to predict soil functional responses to anthropogenic eutrophication and the associated longer-term responses of plant communities to this important global change factor.

openCC (other)Apr 2021View details →
edi52/100

Species Distribution Modeling of Carnivorous Plants Worldwide

Forecasting how carnivorous plant species will respond to climatic change is a key issue in their conservation and management but presents a number of challenges. These challenges derive from interactions between the relatively simplistic statistical methods typically used to forecast species responses to climatic change, which to date have been limited mainly to species distribution models (“SDMs) and particular aspects of the ecology of carnivorous plants, including their rarity, habitat specialization, and limited dispersal ability. The small ranges and oftentimes low local abundance of carnivorous plants provide few occurrence records, which increase the potential for poorly or over-fitted SDMs and misspecification of relationships with their “optimal” environments. The unique habitats in which carnivorous plants often grow also are difficult to characterize using the basic temperature and precipitation data that often undergird SDMs. Rather, habitats in which carnivorous plants are common often are decoupled from broader climatic patterns (e.g., many retain high soil moisture even during seasonal drought) and may be associated with frequent disturbance. Last, dispersal limitation also may constrain range shifts of carnivorous plants as the climate changes. These three issues raise two related questions that are critical for understanding and forecasting the future of carnivorous plants. First, to what extent are current carnivorous plants distributions constrained by climate; and second, how readily, if at all, might carnivorous plants disperse to colonize new habitat as it becomes climatically suitable? We estimated the vulnerability of carnivorous plants to climatic change in light of challenges identified with SDMs in general and their particular application to these unique species. We combined two approaches: “ensembles of small models”, which attempt to deal with the challenges of fitting SDMs for data-limited species; and “bioclimatic velocity”, which is

openCC0Dec 2023View details →
zenodo48/100

Data for "Rapid seaward expansion of seaport footprints worldwide"

<p>[updated November&nbsp;2023]</p><p>This dataset comprises data and code used in "Rapid seaward expansion of seaport footprints worldwide" (Sengupta &amp; Lazarus, 2023; <a href="https://doi.org/10.1038/s43247-023-01110-y">https://doi.org/10.1038/s43247-023-01110-y</a>).</p><p>This repository includes three .csv files, one .xlsx file, and a .ipynb file:</p><p><strong>'Sengupta_Lazarus_REC_1990_2020_v05.csv' </strong>– annual time series of seaward expansion&nbsp;(km^2) between 1990–2020&nbsp;through coastal reclamation&nbsp;for 65 of the world's top 100 container seaports in 2020, as ranked by reported container throughput&nbsp;(<a href="https://lloydslist.maritimeintelligence.informa.com/-/media/lloyds-list/images/top-100-ports-2021/top-100-ports-2021-digital-edition.pdf">Lloyd's List, 2021</a>). Dataset includes Year, Seaport, Country, Region, and Reclaimed area (km^2) [raw measurement].&nbsp;</p><p><strong>'Sengupta_Lazarus_REC_TEU_2011_2020_v03.csv'</strong>&nbsp;–&nbsp;annual time series of seaward expansion&nbsp;(km^2) and reported container throughput (millions TEU)&nbsp;between 2011–2020 for 43 of the world's top 100 container seaports in 2020, as ranked by reported container throughput&nbsp;(<a href="https://lloydslist.maritimeintelligence.informa.com/-/media/lloyds-list/images/top-100-ports-2021/top-100-ports-2021-digital-edition.pdf">Lloyd's List, 2021</a>).&nbsp;Dataset includes Year, Seaport, Country, Region, Reclaimed area (km^2) [raw measurement], and TEU (millions), collated from archived Lloyd's List reports.</p><p><strong>'Sengupta_Lazarus_REC_TEU_totals_v04.csv'</strong> – Simplified dataset listing total seaward expansion&nbsp;(km^2) and container throughput in 2020 for 65 of the world's top 100 container seaports in 2020, as ranked by reported container throughput&nbsp;(<a href="https://lloydslist.maritimeintelligence.informa.com/-/media/lloyds-list/images/top-100-ports-2021/top-100-ports-2021-digital-edition.pdf">Lloyd's List, 2021</a>).&nbsp;Dataset includes Seaport, Country, Region, Reclaimed area (km^2) [raw measurement], ranked list of seaports by expansion extent,&nbsp;and TEU (millions) handled&nbsp;in 2020, and Lloyd's List rank in 2020 (<a href="https://lloydslist.maritimeintelligence.informa.com/-/media/lloyds-list/images/top-100-ports-2021/top-100-ports-2021-digital-edition.pdf">Lloyd's List, 2021</a>).</p><p><strong>'Sengupta_Lazarus_2023_ports_excluded_v2.xlsx'</strong> – contains list of 35 ports excluded from thus analysis because they are either not on an open coastline (e.g. estuarine, riverine) or expanded less than 1 km^2 seaward between 1990–2020.</p><p><strong>'RECLAIM_port_trajectories_v11.ipynb' </strong>– Jupyter notebook for data wrangling and plotting figures presented in Sengupta &amp; Lazarus (2023). (Note that this notebook does not produce the map-based figures presented in that work.)</p><p>The method for calculating reclaimed area over time in Google Earth Engine (GEE) is described in Sengupta et al. (<a href="https://doi.org/10.1029/2022EF002927">2023</a>), and the GEE code is available here:&nbsp;<a href="https://github.com/dhritirajsen/Mapping_Coastal_land_reclamation">https://github.com/dhritirajsen/Seaport_reclamation</a></p><p>These data and code are also available here:&nbsp;<a href="https://github.com/edlazarus/Seaports">https://github.com/edlazarus/Seaports</a></p>

opencc-by-4.0Feb 2023View details →
zenodo48/100

GLOBAL SNAPSHOT Physician Distribution and Density of Physicians per 1000 population - Worldwide 2021

<p>The chart presents the most up-to-date data (2021) available for 49 of the world&acirc;&euro;&trade;s 195 countries, focusing on the total number of physicians and the number of physicians per 1000 population(1). The countries are categorized into four income groups based on World Bank classifications, which are updated annually on July 1st each year(2).</p> <p>Only 25% of the countries present current data. This information is critical for decision-making for healthcare planning and policy development. Equally crucial, is for researchers to have comparable data to propose initiatives, to establish benchmarks and&nbsp; for crafting holistic strategies to gauge and advance progress in healthcare systems globally.</p> <p>Data sources: UnData <a href="https://data.un.org/">https://data.un.org/</a></p> <p>Visualization tools used: RAWGraphs&nbsp;<a href="https://www.rawgraphs.io/">https://www.rawgraphs.io/</a>, MS PowerPoint and Microsoft Excel</p> <p>Intended Audience: Academics and Researchers; Students and Educators; Healthcare Administrators and Policy Makers; Non-Governmental Organizations</p> <p>The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.</p> <p>The NNLM Data Visualization Challenge happens through work funded by the National Institutes of Health's National Library of Medicine, grant number U24LM013751</p> <p>&nbsp;</p> <p>References:</p> <p>1. United Nations, Department of Economic and Social Affairs. 10 Health Personnel. In: Statistical Yearbook. 66th issue (2023). New York: United Nations; 2023. (ST/ESA/STAT/SER.S/42). [Dataset available at UnData] <a href="https://data.un.org/_Docs/SYB/CSV/SYB66_154_202310_Health%20Personnel.csv">https://data.un.org/_Docs/SYB/CSV/SYB66_154_202310_Health%20Personnel.csv</a></p> <p>2 World Bank. World Bank Country and Lending Groups. World Bank Data Help Desk [Internet]. [cited 2024 Apr 5]. Available from:<a href="https://datahelpdesk.worldbank.org/knowledgebase/articles/906519-world-bank-country-and-lending-groups"> https://datahelpdesk.worldbank.org/knowledgebase/articles/906519-world-bank-country-and-lending-groups</a></p>

opencc-by-4.0May 2024View details →
zenodo48/100

Data from: Flock size and structure influence reproductive success in four species of flamingo in 540 captive populations worldwide

<p><strong>Summary</strong></p> <p>This dataset accompanies the publication &quot;<strong>Flock size and structure influence reproductive success in four species of flamingo in 540 captive populations worldwide</strong>&quot; published in Zoo Biology. It contains anonymised data from 540 captive flamingo populations, and includes the four species:&nbsp;<em>Phoeniconaias minor, Phoenicopterus chilensis, Phoenicopterus roseus</em> and<em> Phoenicopterus ruber</em>.&nbsp;Data were sourced from the&nbsp;Zoological Information Management System (ZIMS), operated by Species360 (https://www.species360.org/). ZIMS is the largest real-time database of comprehensive and standardized information spanning more than 1,200 zoological collections globally, and provides the number of institutions currently managing each flamingo species and both their current and historic population sizes.&nbsp;These data were used to&nbsp;investigate the relationship between reproductive success and both flock size, and structure, on a global scale.</p> <p>This dataset also contains climatic data&nbsp;provided by WorldClim, which were used to assess&nbsp;the influence of climatic variables on captive flamingo reproductive success globally. The WorldClim database averages 19 different climatic variables derived from monthly temperature and rainfall values at a 1 km spatial resolution for the period 1970-2000. Using geographic coordinates (latitude and longitude) we calculated several climatic metrics for each institution.&nbsp;</p> <p>&nbsp;</p> <p><strong>Description of the Dataset</strong></p> <p>One file is provided for each species (<em>P. minor, P. chilensis, P. roseus </em>and&nbsp;<em>P. ruber</em>)&nbsp;as a csv file. Each file contains the following 15 columns:</p> <ul> <li><strong>Institution Code: </strong>An anonymous code used to identify individual zoological institutions.&nbsp; &nbsp; &nbsp; &nbsp;</li> <li><strong>Country: </strong>The country where the institution is located.</li> <li><strong>Year: </strong>Current year (<em>t</em>).</li> <li><strong>Flock Size:</strong> Flock size in year <em>t.</em></li> <li><strong>Males: </strong>The number of males in the flock in year <em>t.</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</li> <li><strong>Females:</strong> The number of females in the flock in year <em>t.</em></li> <li><strong>Unsexed:</strong> The number of unsexed individuals in the flock in year <em>t.</em></li> <li><strong>Proportion of Females: </strong>The proportion of the flock made up of female individuals in year <em>t</em>.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</li> <li><strong>Proportion of Unsexed:</strong> The proportion of the flock made up of unsexed individuals in year <em>t.</em></li> <li><strong>Hatches:</strong> Number of birds hatched in year <em>t.</em></li> <li><strong>Proportion of Additions:</strong> The proportion of the flock in year <em>t</em> made up of additions from year <em>t-1</em> (not including new birds hatched into the flock).</li> <li><strong>MAP: </strong>Mean annual precipitation (mm).</li> <li><strong>MAT: </strong>Mean annual temperature (&deg;C).</li> <li><strong>MAP Var: </strong>Mean annual variation in precipitation (MAP coefficient of variation).</li> <li><strong>MAT Var: </strong>Mean annual variation in temperature (MAT standard deviation).</li> </ul> <p>Note: Mean Annual Temperature (MAT) is provided by WorldClim as &deg;C multiplied by 10, and similarly mean annual variation in temperature as MAT standard deviation multiplied by 100. In the corresponding publication, both were divided (by 10 and 100 respectively) prior to modelling to avoid confusion in the units used.</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>We acknowledge and thank all Species360 member institutions for their continued support and data input. The research which data refers to was funded by the Irish Research Council Laureate Awards 2017/2018 IRCLA/2017/60 to Y.M.B. Additionally, S.Q.S. received funding from the International Max Planck Research School for Organismal Biology. The Species360 Conservation Science Alliance would like to thank their sponsors: the World Association of Zoos and Aquariums, Wildlife Reserves of Singapore, and Copenhagen Zoo.&nbsp;</p> <p>&nbsp;</p> <p><strong>Disclaimer</strong></p> <p>Despite our best efforts at screening the data for errors and inconsistencies, some information could be erroneous. Similarly, data contained within&nbsp;ZIMS are based on submitted records from individual institutions, and are not&nbsp;subject&nbsp;to editorial verification, potentially permitting errors or failure to update species holdings etc. Despite this, ZIMS represents the only global database&nbsp;of zoo collection composition records, and as a result,&nbsp;is used by the IUCN, Convention on International Trade in Endangered Species (CITES), the Wildlife Trade Monitoring Network (TRAFFIC), United States Fish and Wildlife Service (USFWS) and Department for Environment, Food and Rural Affairs (DEFRA).&nbsp;</p> <p>&nbsp;</p> <p><strong>Credit</strong></p> <p>If you use this dataset, please cite the corresponding publication:</p> <p>Mooney, A., Teare, J. A., Staerk, J.,Smeele, S. Q., Rose, P., Edell, R. H., King, C. E., Conrad, L., &amp; Buckley, Y. M. (2023). Flock size and structure influence reproductive success in four species of flamingo in 540 captive populations worldwide.<em> Zoo Biology</em>, 1&ndash;14. <a href="https://doi.org/10.1002/zoo.21753">https://doi.org/10.1002/zoo.21753</a></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
edi48/100

CoRRE Trait Data: A collection of 17 categorical and continuous traits for more than 4000 grassland species worldwide

In our changing world, it is critical to understand and predict plant community responses to global change drivers. Plant functional traits promise to be a key predictive tool for many ecosystems, including grasslands, however their use requires both complete plant community and functional trait data. Yet, representation of these data in global databases is incredibly sparse, particularly beyond a handful of most used traits and common species. Here we present the CoRRE Trait Database, spanning 17 traits (9 categorical, 8 continuous) anticipated to predict species’ responses to global change for 4,079 vascular plant species across 173 plant families present in 390 grassland experiments from around the world. The database contains complete categorical trait records for all 4,079 plant species, obtained from a comprehensive literature search. Additionally, the database contains nearly complete coverage (99.97%) of species mean values for continuous traits for a subset of 2,927 plant species, predicted from observed trait data drawn from TRY and a variety of other plant trait databases using Bayesian Probabilistic Matrix Factorization (BHPMF) and multivariate imputation using chained equations (MICE). These data will shed light on mechanisms underlying population, community, and ecosystem responses to global change in grasslands worldwide.

openCC BYMay 2024View details →
zenodo44/100

African Swine Fever Worldwide Epidemiology Data - OIE Webscrape example - Geocoded using Google API and Manual

<p>Example African Swine Fever dataset generated by programs described in following publication&nbsp;</p> <p>Title: Web-scraping programmatic techniques in aggregating difficult to access OIE WAHIS animal disease outbreak information; using African Swine Fever in Europe as an example.</p> <p>Short running title: Methods for web-scraping OIE WAHIS data.</p> <p>Abstract: This study describes and makes available new methods for acquiring difficult to access, publicly available, disease surveillance data. It uses World Organisation for Animal Heath (OIE) data on African Swine Fever (ASF) outbreaks in Belarus and its neighbouring European countries to showcase the importance of adequate disease surveillance data to inform decision-making. The data acquired from these methods allow for large-scale, geospatial outbreak mapping and summary statistics of any terrestrial disease listed on the OIE World Animal Health Information System (WAHIS) database. These techniques will make important epidemiological data more accessible to the scientific community and aid in gaining further insight into the occurrence and spread of OIE listed diseases in a timely manner, fulfilling an important function of disease surveillance.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Worldwide Soundscapes project metadata and analysis scripts

<p>The Worldwide Soundscapes project is a global, open inventory of spatio-temporally replicated passive acoustic monitoring meta-datasets (i.e. meta-data collections). This Zenodo entry comprises the data tables that constitute its (meta-)database, as well as their description. Additionally, R scripts are provided to replicate the analysis published in [placeholder].</p> <p>The overview of all sampling sites and timelines can be found on the corresponding project on&nbsp;<a href="https://ecosound-web.de/ecosound_web/collection/index/106">ecoSound-web</a>, as well as a <a href="https://ecosound-web.de/ecosound_web/collection/show/49">demonstration collection</a> containing selected recordings. The recordings of this collection were annotated and analysed to explore macro-ecological trends.</p> <p>The&nbsp;audio recording criteria justifying inclusion into the meta-database are:</p> <ul> <li>Stationary (no transects, towed sensors or microphones mounted on cars)</li> <li>Passive (unattended, no human disturbance by the recordist)</li> <li>Ambient (no directional microphone or triggered recordings, non-experimental conditions)</li> <li>Spatially and/or temporally replicated (i.e. multiple sites sampled at the same time and/or multiple days - covering the same daytime - sampled at the same site)</li> </ul> <p>The individual columns of the provided data tables are described in the following. Data tables are linked through primary keys; joining them will result in a database. The data shared here only includes validated collections.</p> <p><strong>Changes from version 4.0.0</strong></p> <p>Added link to the published synthesis.</p> <p><strong>Meta-database CSV files</strong></p> <p><strong>collections</strong></p> <ul> <li>collection_id: unique integer, primary key</li> <li>name: name of the dataset. if it is repeated,&nbsp; incremental integers should be used in the "subset" column to differentiate them.</li> <li>ecoSound-web_link: link of validated meta-collection on ecoSound-web</li> <li>primary_contributors: full names of people deemed corresponding contributors who are responsible for the dataset</li> <li>secondary_contributors: full names of people who are not primary contributors but who have significantly contributed to the dataset, and who could be contacted for in-depth analyses</li> <li>date_added: when the datased was added (YYYY-MM-DD)</li> <li>URL_open_recordings: internet link for openly-available recordings from this collection</li> <li>URL_project: internet link for further information about the corresponding project</li> <li>DOI_publication: Digital Object Identifiers of corresponding publications</li> <li>core_realm_IUCN: The main, core realm of the dataset according to IUCN Global Ecosystem Typology (v2.0): https://global-ecosystems.org/</li> <li>medium: the physical medium the microphone is situated in</li> <li>locality: optional free text about the locality</li> <li>contributor_comments: free-text field for comments by the primary contributors</li> </ul> <p><strong>collections-sites</strong></p> <ul> <li>dataset_ID: primary key of collections table</li> <li>site_ID: primary key of sites table</li> </ul> <p><strong>sites</strong></p> <ul> <li>site_ID: unique integer, primary key</li> <li>site_name: internal name or code of sampling site as used in respective projects</li> <li>latitude_numeric: site's numeric degrees of latitude</li> <li>longitude_numeric: site's numeric degrees of longitude</li> <li>blurred_coordinates: whether latitude and longitude coordinates are inaccurate, boolean. Coordinates may be blurred with random offsets, rounding, snapping, etc. Indicate the blurring method inside the comments field</li> <li>topography_m: vertical position of the microphone relative to the sea level. for sites on land: elevation. For marine sites: depth (negative). in meters. Only indicate if the values were measured by the collaborator.</li> <li>freshwater_depth_m: microphone depth, only used for sites inside freshwater bodies that also have an elevation value above the sea level</li> <li>realm: Ecosystem type: main realm according to IUCN GET&nbsp; https://global-ecosystems.org/</li> <li>biome: Ecosystem type: main biome according to IUCN GET&nbsp; https://global-ecosystems.org/</li> <li>functional_group: Ecosystem type: main functional group according to IUCN GET &nbsp;https://global-ecosystems.org/</li> <li>contributor_comments: free text field for contributor comments</li> <li>GADM_0: Global ADMinistrative Database level 0 classification of terrestrial site or marine site that is within territorial waters. Source: https://gadm.org/download_world.html</li> <li>IHO: International Hydrographic Organization classification of marine site. Source: https://marineregions.org/downloads.php</li> <li>WDPA: World Database on Protected Areas classification of the site. Source: https://www.protectedplanet.net/en/thematic-areas/wdpa?tab=WDPA</li> </ul> <p><strong>deployments</strong></p> <ul> <li>dataset_ID: primary key of datasets table</li> <li>deployment: identical subscript letters to denote rows that belong to the same deployment. For instance, you may use different operation times and schedules for different target taxa within one deployment.</li> <li>subset_site_ID: If the deployment was not done in all the sites of the corresponding collection, site IDs where the deployment was conducted</li> <li>start_date: date of deployment start</li> <li>start_time_mixed: deployment start local time, either in HH:MM format or a choice of solar daytimes (sunrise, sunset). Corresponds to the recording start time for continuous recording deployments. If multiple start times were used, you should mention the latest start time (corresponds to the earliest daytime from which all recorders are active).&nbsp;If applicable, positive or negative offsets from solar times can be mentioned (For example: if data are collected one hour before sunrise, this will be "sunrise-60")</li> <li>permanent: whether the deployment is permanent, boolean</li> <li>end_date: date of deployment end (date when last scheduled operation starts)</li> <li>end_time_mixed: deployment end local time, either in HH:MM format or a choice of solar daytimes (sunrise, sunset, noon, midnight). Corresponds to the recording end time for continuous recording deployments.</li> <li>operation_mode: continuous: recording takes place from the deployment start date-time to deployment end date-time.<br>periodical: recording takes place periodically (i.e., with duty cycle) from the deployment start date-time to deployment end date-time.<br>scheduled: recording takes place &nbsp;during scheduled daily time intervals (optionally with duty cycle)</li> <li>duty_cycle_minutes: duty cycle of the recording (i.e. the fraction of minutes when it is recording), written as "recording(minutes)/period(minutes)". empty if no duty cycle is used.&nbsp;For example: "1/6" if the recorder is active for 1 minute and standing by for 5 minutes</li> <li>operation_start_time_mixed: only for scheduled recordings: start local time, either in HH:MM format or a choice of solar daytimes (sunrise, sunset, noon, midnight).&nbsp;If applicable, positive or negative offsets from solar times can be mentioned (For example: if data are collected one hour before sunrise, this will be "sunrise-60")</li> <li>operation_duration_minutes: only for scheduled recordings: duration of operation in minutes, if constant</li> <li>operation_end_time_mixed: only for scheduled recordings: end local time, either in HH:MM format or a choice of solar daytimes (sunrise, sunset, noon, midnight). Only required if durations are variable. Do not use when end times are ambiguous (for instance, if a recording could be 1 hour or 25 hours long because the end is on the next day).&nbsp;If applicable, positive or negative offsets from solar times can be mentioned (For example: if data are collected one hour before sunrise, this will be "sunrise-60")</li> <li>high_pass_filter_Hz: frequency of the high-pass filter of the recorder if applied, in Hz. Otherwise, write "none". This may be called a "low-cut" filter too.</li> <li>bit_depth: sampling bit depth of the recordings. Often constant for a particular recorder</li> <li>channels: number of recorded audio channels</li> <li>sampling_frequency_kHz: frequency at which the microphone signal was sampled by the recorder&nbsp;(sounds of half that frequency will be recorded)</li> <li>recorder: recorder used for deployment</li> <li>microphone: microphone used for deployment</li> <li>target_taxa: main IUCN animal taxa that were studied with this deployment, using the exact IUCN Red list names (http://www.iucnredlist.org/), separated by commas. Only genera, families, orders, and classes are accepted. Empty if there was no taxonomic focus (i.e., general soundscapes were the study focus).</li> <li>contributor_comments: free text field for contributor comments</li> <li>exact_recordings: whether the deployment data here have been superseded by inserting more exact recording date-time ranges into the meta-collection on ecoSound-web</li> </ul> <p><strong>recordings (partial download from <a href="https://ecosound-web.de/">ecoSound-web</a>)</strong></p> <ul> <li>recording_id: primary key of the recordings table</li> <li>collection_id: ID of the collection the recording belongs to</li> <li>name: name of the recording</li> <li>site_id: site ID the recording belongs to:</li> <li>recorder_id: ID of the recorder used for the recording (internal ecoSound-web code)</li> <li>microphone_id: ID of the microphone used for the recording (internal ecoSound-web code)</li> <li>recording_gain:recording gain applied for amplifying the audio signal, in decibels</li> <li>duty_cycle_recording: fraction of the recording periode when the recorder is actively recording audio</li> <li>duty_cycle_period: period of the duty cycle, i.e., time between the starts of two subsequent recordings</li> <li>note: comments (contains the target taxon)</li> <li>file_date: date of the recording start</li> <li>file_time: local time of the recording start</li> <li>sampling_rate: audio sampling rate in Hz</li> <li>bitdepth: depth in bits for each audio sample</li> <li>channel_num: number of channels</li> <li>duration: duration of the recording in seconds. Note: duty-cycled recordings cover only a proportion of this duration<strong><br></strong></li> </ul> <p><strong>affiliations</strong></p> <ul> <li>affiliation_id: primary key of affiliations table</li> <li>lab_research_group: Laboratory or research group name</li> <li>department_school_institute: department, school, or institute name</li> <li>university_institution: University or institution name</li> <li>street_address: street address</li> <li>region_state_province_city: region, state, province, or city name</li> <li>postal_code: postal code</li> <li>country: country name</li> </ul> <p><strong>primary_contributors</strong></p> <ul> <li>First_name: First, given name, anonymised when contributor is technically accepted but has not yet given publication authorisation</li> <li>Last_name: Last, family name, anonymised when contributor is technically accepted but has not yet given publication authorisation</li> <li>ORCiD</li> <li>affiliation_IDs: primary keys of the affiliations' table corresponding affiliations, separated by comma</li> <li>first_tier_position: Author position in first-tier</li> <li>publication_agreement: Has contributor explicitly agreed to share her/his meta-data in the collaboration agreement?</li> <li>co_author_first_synthesis: Has contributor confirmed co-authorship intention in the collaboration agreement?</li> </ul> <p>The following columns describe the contributor's role in the project accordint to <a href="https://credit.niso.org/">CRediT</a> taxonomy.</p> <p><strong>Auxiliary files for reproducing analysis</strong></p> <p><strong>R scripts</strong></p> <ul> <li><strong>acoustic analysis.R:&nbsp;</strong>reproduces the result of the soundscape case studies</li> <li><strong>metadata analysis.R:</strong> reproduces the metadata analysis results in the publication</li> </ul> <p><strong>Data from the demonstration collection (download from ecoSound-web)</strong></p> <ul> <li><strong>demo_recordings.csv:</strong> metadata of the recordings, see recordings table</li> <li><strong>demo_sites.csv: </strong>metadata of the sampling locations, see sites table</li> <li><strong>demo_tags.csv: </strong>data describing annotations made in demonstration recordings for the biophony, anthropophony, geophony, and unknown sound sources</li> <li><strong>spectrograms.zip:</strong> contains PNG format spectrograms used in generating Figure 5</li> </ul> <p><strong>Externally sourced data</strong></p> <ul> <li><strong>GET_areas_2.1.1.csv:&nbsp;</strong>raw data obtained from Keith et al. 2023 (https://doi.org/10.5281/zenodo.10081251), then summarized in QGIS to obtain areas per functional group</li> <li><strong>Havlik_sites.csv:</strong> data obtained from Havlik et al. 2022 supplementary material (https://www.frontiersin.org/articles/10.3389/fmars.2022.919418), originally named "Data Sheet 1.CSV"</li> <li><strong>Sugai_sites_updated.csv:</strong> data obtained from Sugai et al. 2019 (https://doi.org/10.1093/biosci/biy147), personal communication with permission</li> <li><strong>taxonomy.csv:</strong> raw data obtained from IUCN Red List for all animal taxa (https://www.iucnredlist.org/)</li> <li><strong>topography_range_latitude.csv:</strong> raw topography from GEBCO sub-ice data (https://www.gebco.net/data_and_products/gridded_bathymetry_data/), summarised by bins of 10 latitudinal rows</li> </ul>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Covid-19 Worldwide Data 2021 Cases and Vaccination

<p>This dataset includes information about active cases, accumulative cases, accumulative deaths, daily information, vaccination classified by country of year 2021.&nbsp;&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

AMPSphere : the worldwide survey of prokaryotic antimicrobial peptides

<p><strong>AMPSphere v.2022-03: the worldwide survey of prokaryotic antimicrobial peptides</strong></p> <p>&nbsp;</p> <p><strong>INTRODUCTION</strong></p> <p>AMPSphere is a comprehensive catalog of antimicrobial peptides predicted using Macrel (DOI: <a href="https://peerj.com/articles/10555/">10.7717/peerj.10555</a>) from 63,410 public metagenomes, <a href="http://progenomes.embl.de/">ProGenomes v2.2 database</a> (82,400 high-quality microbial genomes), and c.a. 4k non-whitelisted microbial genomes from NCBI. Currently, AMPSphere is available as a web resource at <a href="https://ampsphere.big-data-biology.org/">https://ampsphere.big-data-biology.org/</a>.</p> <p>&nbsp;</p> <p><strong>GENERATION</strong></p> <p>Peptides were predicted using <a href="https://www.big-data-biology.org/software/macrel/">Macrel</a>. Singleton peptides were removed, except those with a direct hit to <a href="http://dramp.cpu-bioinfor.org/">DRAMP<br> database</a>. Redundant peptides were coded using a reduced alphabet and hierarchically clustered using CD-HIT (version 4.6) at 100%, 85%, and 75% of amino acid identity (and 90% of overlap of the shorter peptide). The obtained clusters were numbered by decreasing size (number of peptides). Each level of clustering was called a SPHERE. Redundant nucleotide sequences for the gene variants of different AMPs also were included in this version of AMPSphere.</p> <p>&nbsp;</p> <p><strong>STATISTICS</strong></p> <p>AMPSphere v.2022-03 contains 863,498 sequences (avg length: 36 amino acids, range 8-98). DRAMP database was used to find confirmed sequences with strict homology to reference. This approach showed that 2,488 peptides were previously confirmed in our dataset.</p> <p>&nbsp;</p> <p><strong>IDENTIFIERS</strong></p> <p>Peptides are named in the form <strong>&gt;AMP10.XXX_XXX</strong> where <strong>XXX_XXX</strong> is a unique numerical identifier (starting at zero). Numbers were assigned in order of increasing number of copies. So that the lower the number, the greater number of copies of that peptide were present in the input data. Annotations were also provided as separated fields in the fasta file, containing their:</p> <p>- SPHERE families at level III (corresponding to hierarchically obtained clusters using 100-85-75% of identity with a minimum overlap of 90% of the shorter gene).</p> <p>Example of header:</p> <pre><code class="language-bash">&gt;AMP10.000_000 | SPHERE-III.001_493 </code></pre> <p><br> <strong>VERSION DETAILS</strong></p> <p><strong>Version 2022-03</strong> includes:</p> <p>- quality assessment of documented AMPs,</p> <p>- metadata associated with the genes,</p> <p>- a better taxonomic identification of AMP sources using GTDB.</p> <p><em>WARNING: Due to a different procedure of AMP sorting, now some entries and families may have changed their accessions.</em></p> <p>&nbsp;</p> <p><strong>FILES WITHIN THIS VERSION</strong></p> <p><em>README.md</em><br> This file.</p> <p>&nbsp;</p> <p><em>AMPsphere_v.2022-03.fna.xz</em><br> Multi-fasta with AMPSphere gene sequences (nucleotide).</p> <p>&nbsp;</p> <p><em>AMPsphere_v.2022-03.faa.gz</em></p> <p>Multi-fasta with AMPSphere peptide sequences (amino acid).</p> <p>&nbsp;</p> <p><em>SPHERE_v.2022-03.levels_assessment.tsv.gz</em></p> <p>TSV table relating AMP name and the hierarchically obtained clusters per level. Columns:</p> <p>- AMP accession<br> - evaluation vs. representative<br> - SPHERE_fam level I<br> - SPHERE_fam level II<br> - SPHERE_fam level III</p> <p>Levels of each SPHERE family:</p> <p>I: contains clusters obtained with 100% of identity cut-off and 90% of overlap of the shorter sequence;</p> <p>II: contains clusters obtained with the unclustered sequences and the representatives from level I at 85% of identity and 90% of overlap of the sorter sequence;</p> <p>III: contains clusters obtained with the unclustered sequences and the representatives from level II at 75% of identity and 90% of overlap of the sorter sequence;</p> <p>`evaluation vs. representatives` shows the percent of identity the sequence has in an alignment against the cluster representative, and also the overlap in percent.</p> <p>Example:</p> <p>&nbsp;&nbsp;&nbsp; * -- This means: this sequence is a cluster representative.</p> <p>OR something like this:</p> <p>&nbsp;&nbsp;&nbsp; 77.50%,1:40:1:40 -- This means: alignment identity against the<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; representative of the cluster equals 77.5% and the<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; alignment start and end position for the query (1 and<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 40, respectively), and target (1 and 40, respectively).</p> <p>&nbsp;</p> <p><em>AMPSphere_v.2022-03.quality_assessment.tsv.gz</em><br> TSV table containing the results of each quality test (by sequence). Columns:</p> <p>- AMP ID<br> - Antifam<br> - RNAcode<br> - Metaproteomes<br> - Metatranscriptomes<br> - Coordinates</p> <p>Results are one of &#39;Passed&#39;, &#39;Failed&#39;, or &#39;Not tested&#39;.</p> <p><a href="https://www.ebi.ac.uk/research/bateman/software/antifam-tool-identify-spurious-proteins">Antifam </a>results show if the sequence matches (&#39;Fail&#39;) or does not match (&#39;Pass&#39;) to Antifams, a set of well-known spurious ORFs.</p> <p><a href="https://github.com/ViennaRNA/RNAcode">RNAcode</a> relies on gene diversity, therefore, families with less than 3 different gene sequences could not be tested and were marked as such.</p> <p>The direct match of 50% of our peptide to transcripts (in at least 2 different samples) or peptides from meta-omics studies sampled from different environments assigned the peptide as passing the metatranscriptomes and metaproteomes tests, respectively.</p> <p>Finally, the coordinates test check if the start of the small ORF happens with at least one stop codon upstream, this ensures that the gene is not a fragment from a larger protein.</p> <p>&nbsp;</p> <p><em>AMPSphere_v.2022-03.general_geneinfo.tsv.gz</em><br> TSV table relating AMP, gene name, the microbial source, sample, environment, and geographical location. Columns:</p> <p>- gmsc (gene code access)&nbsp;&nbsp; &nbsp;<br> - amp&nbsp;&nbsp; &nbsp;<br> - sample (biosample)<br> - source (microbial origin, GTDB taxonomy)<br> - specI (species cluster according to ProGenomes v.2 classification)<br> - is_metagenomic (False if comming from a high-quality microbial genome)<br> - geographic_location<br> - latitude<br> - longitude<br> - general_envo_name<br> - environment_material</p> <p>&nbsp;</p> <p><strong>CONTACT</strong></p> <p>You can contact us via our discussion group: <a href="https://groups.google.com/g/ampsphere-users">https://groups.google.com/g/ampsphere-users</a></p> <p>AMPsphere main developers:</p> <p>- <a href="mailto:celio@big-data-biology.com?subject=AMPSphere%20v.2022-03&amp;body=Dear%20Celio%2C%20%0A%0ARegarding%20AMPSphere%20v.2022-03.">C&eacute;lio Dias Santos J&uacute;nior</a><br> - <a href="mailto:yiqian@big-data-biology.org?subject=AMPSphere%20v2022-03&amp;body=Dear%20Yiqian%2C%20%0A%0ARegarding%20AMPSphere%20v2022-03.%0A">Yiqian Duan</a><br> - <a href="mailto:hui@big-data-biology.org?subject=AMPSphere%20v2022-03&amp;body=Dear%20Hui%2C%20%0A%0ARegarding%20AMPSphere%20v.2022-03.">Hui Chong</a><br> - <a href="mailto:luispedro@big-data-biology.com?subject=AMPSphere%20v.2022-03&amp;body=Dear%20Luis%2C%20%0A%0ARegarding%20AMPSphere%20v.2022-03.">Luis Pedro Coelho</a></p> <p><br> <strong>COPYRIGHT NOTICE</strong></p> <p><em>AMPSphere v.2022-03 - the worldwide survey of prokaryotic antimicrobial peptides.</em></p> <p>This work is a joint effort of Big Data Biology group from the Institute of Science and Technology for Brain-Inspired Intelligence (ISTBI) - Fudan University, Shanghai, China, and the Structural and Computational Biology Unit<br> (Heidelberg) - European Molecular Biology Laboratory (EMBL).</p> <p>Copyright (C) 2019-2022 The Authors</p> <p>&nbsp;&nbsp; AMPSphere IS PROVIDED &quot;AS IS&quot;, WITHOUT WARRANTY OF ANY KIND,<br> &nbsp;&nbsp; EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES<br> &nbsp;&nbsp; OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.<br> &nbsp;&nbsp; IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM,<br> &nbsp;&nbsp; DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR<br> &nbsp;&nbsp; OTHERWISE,ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE<br> &nbsp;&nbsp; USE OR OTHER DEALINGS IN THE SOFTWARE.</p> <p>&nbsp;&nbsp; This database is free; you can redistribute it and/or modify it<br> &nbsp;&nbsp; as you wish, under the terms of the CC BY 4.0 license.</p> <p>&nbsp;&nbsp; You are allowed to:</p> <p>&nbsp;&nbsp; Share &mdash; copy and redistribute the material in any medium or format</p> <p>&nbsp;&nbsp; Adapt &mdash; remix, transform, and build upon the material for any purpose,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; even commercially.</p> <p>&nbsp;&nbsp; You may also obtain a copy of the CC BY 4.0 license here:<br> &nbsp; &nbsp;<br> &nbsp;&nbsp; https://creativecommons.org/licenses/by/4.0/</p> <p><br> <strong>REFERENCES CITED</strong></p> <p>- Macrel: Santos-J&uacute;nior CD, Pan S, Zhao X, Coelho LP. 2020. Macrel: antimicrobial peptide screening in genomes and metagenomes. PeerJ 8:e10555. https://doi.org/10.7717/peerj.10555</p> <p>- ProGenomes: Mende DR, Letunic I, Maistrenko OM et al. 2020. proGenomes2: an improved database for accurate and consistent habitat, taxonomic and functional annotations of prokaryotic genomes.&nbsp; Nucleic Acids Research 48(D1): D621&ndash;D625. https://doi.org/10.1093/nar/gkz1002</p> <p>- DRAMP: Kang X, Dong F, Shi C et al. 2019. DRAMP 2.0, an updated data repository of antimicrobial peptides. Sci Data 6, 148. https://doi.org/10.1038/s41597-019-0154-y</p> <p>- ANTIFAM: Eberhardt RY, Haft DH, Punta M, Martin M, O&rsquo;Donovan C, BatemanA. 2012. AntiFam: a tool to help identify spurious ORFs in protein annotation. Database, Bas003.</p> <p>- RNAcode: Washietl S, Findeiss S, M&uuml;ller SA, Kalkhof S, von Bergen M, Hofacker IL, Stadler PF, Goldman N. 2011. RNAcode: robust discrimination of coding and noncoding regions in comparative sequence data. RNA 17(4):578-94.<br> &nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Supplementary material: Negative valence in Obsessive-Compulsive Disorder: A worldwide mega-analysis of task-based functional neuroimaging data of the ENIGMA-OCD consortium

<p>The ridge plots attached here accompany the supplement to the manuscript <em>Negative valence in Obsessive-Compulsive Disorder: A worldwide mega-analysis of task-based functional neuroimaging data of the ENIGMA-OCD consortium</em>&nbsp;(Dzinalija et al., 2024)<em>. </em>These are full results of Figures 1, 2C, and 3C of the manuscript and Figures S3, S5, S7 and S9 of the supplement depicting whole-brain Bayesian multilevel models run using the Regional Bayesian Analysis toolbox (RBA; Chen et al., 2019). Whole-brain analyses were parcelated into the Schaefer-Yeo 7-network 200-parcel cortical atlas (Schaefer et al., 2018) and the Melbourne 32-region subcortical atlas (Tian et al., 2020). Results are presented according to contrast of interests: [Negative &gt; Neutral], [OCD &gt; Neutral], [Threat &gt; Neutral], and [OCD &gt; Threat], and effects of interest: [Diagnosis = OCD or HC], [MED = medication], [AO = age of onset], [YBOCS = OCD severity], and [Intercept = task effect].&nbsp;</p> <p>P+ values denote the probability that there is increased brain activation in a given region of the Schaefer 200-parcel 7-network cortical atlas and Melbourne 32-region subcortical atlas. We used the guidelines proposed by Chen et al. (2019) to infer credibility of evidence, namely taking a positive posterior probability (P+) of &lt;0.10 or &gt;0.90 as indication of moderate evidence and, &lt;0.05 or &gt;0.95 or &lt;0.025 or &gt;0.975 as strong or very strong evidence, respectively. To interpret the pairwise comparisons presented as Group1-vs-Group2, posterior distributions to the right of the green no-effect line represent regions in which individuals in Group 1 show credible evidence for higher activation than individuals in Group 2. Regions with posterior distributions to the left of this line show credible evidence for higher activity in Group 2 than in Group 1.&nbsp; (Darker) red color represents regions in which individuals in Group 1 show moderate-to-very-strong evidence for higher activation than Group 2. (Darker) blue color represents regions in which Group 2 show moderate-to-very-strong evidence for higher activation than Group 1. Grey color represents regions in which there is no strong evidence of a difference between Group 1 and Group 2. Schaefer-Yeo 200-parcel atlas name abbreviations can be retrieved <a href="https://github.com/ThomasYeoLab/CBIG/tree/master/stable_projects/brain_parcellation/Schaefer2018_LocalGlobal/Parcellations">via this link.</a></p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Datasample - Speed Performance of 121 LAMs webpages worldwide

<p>This particular dataset is composed of 121 cases of homepages speed performance coming from Libraries, Archives and Museums worldwide. Data for 12 metrics are taking place. 6 for mobile and six for desktop. PageSpeed Insights of Google has been used to retrieved data.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

A data directory to facilitate investigations on worldwide wildlife trafficking

<p>We describe a novel, open-access data directory &nbsp;on wildlife trafficking and a corresponding visualization tool that can be used to identify data for multiple purposes, such as exploring wildlife trafficking hotspots and convergence points with other crime, discovering key drivers or deterrents of wildlife trafficking, and uncovering structural patterns. Keyword searches, expert elicitation, and peer-reviewed publications were used to search for extant sources used by industry and non-profit organizations, as well as those leveraged to publish academic research articles.&nbsp; The open-access data directory is designed to be a living document and searchable according to multiple measures. The directory can be instrumental in the data-driven analysis of unsustainable illegal wildlife trade, supply chain structure via link prediction models, the value of demand and supply reduction initiatives via multi-item knapsack problems, or trafficking behavior and transportation choices via network interdiction problems.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Worldwide CO2 emissions and natural disasters from 1960 to 2021

<p>A CSV file containing worldwide CO2 emissions as well as the number of natural disasters per year.</p> <p>Sources:</p> <ul> <li>Global Carbon Atlas <ul> <li>DOI:&nbsp;<a href="http://doi.org/10.17616/R3434K">http://doi.org/10.17616/R3434K</a></li> <li>URL:&nbsp;<a href="http://www.globalcarbonatlas.org/en/CO2-emissions">http://www.globalcarbonatlas.org/en/CO2-emissions</a></li> <li>Last accessed: 2023-05-09</li> </ul> </li> <li>EM-DAT <ul> <li>DOI:&nbsp;<a href="http://doi.org/10.17616/R3QQ1X">http://doi.org/10.17616/R3QQ1X</a></li> <li>URL:&nbsp;<a href="https://public.emdat.be/data">https://public.emdat.be/data</a>&nbsp;(registration necessary)</li> <li>Last accessed: 2023-05-14</li> </ul> </li> <li>GitHub Project <ul> <li>DOI: <a href="http://doi.org/10.5281/zenodo.7934702">http://doi.org/10.5281/zenodo.7934702</a>&nbsp;</li> <li>URL:&nbsp;<a href="https://github.com/jkopec/global-emission-and-disaster-analysis">https://github.com/jkopec/global-emission-and-disaster-analysis</a></li> </ul> </li> </ul>

opencc-by-4.0May 2023View details →
zenodo44/100

Dataset Worldwide Survey on the Impact of AI Chatbots and Large Language Models in Dental Education: Insights from Dental Educators

<p><strong>This dataset contains responses from participants regarding their awareness, knowledge, and perceptions of AI-powered tools in dental education. The data was collected during May-June 2023 to investigate the potential enhancement that AI can bring to dental education. The dataset includes variables related to participants&#39; demographics, experiences, perceptions, and opinions.</strong></p> <p><strong>Details in the published protocol by Uribe, S. E., &amp; Maldupa, I. (2023, June 2). Chatbots In Dental Education - Research Protocol. https://doi.org/10.17605/OSF.IO/3BSG2</strong></p>

opencc-by-4.0Jun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record