Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
225
datasets available to search
ShareScore release 0.9.0
Dataset results
225 results for “Global database”
Global sea turtle epibiont database
Open the record for dataset details and reuse information.
A cross-checked global monthly weather station database for precipitation covering the period 1901 to 2010
<p>This database entry represents a comprehensive compilation of monthly weather station records for precipitation from multiple data sources for the period 1901-2010, with an emphasis on climate normal averages for the period 1961-1990. The database corresponds to the journal publication: Castellanos-Acuña, D. and Hamann, A. 2020. A cross-checked global monthly weather station database for precipitation covering the period 1901 to 2010. Geoscience Data Journal (https://rmets.onlinelibrary.wiley.com/journal/20496060, article in press, January 2020).</p> <p>We use digital elevation models and nearby stations to search for inconsistencies in reported station locations and recorded precipitation values. We also estimated missing values in weather station time series using a linear model approach based on interpolated anomaly surfaces. The resulting station records were ranked into ten classes, according to the completeness of records, the reliability of missing value estimations and other criteria. We corrected incomplete or erroneous location and elevation information for 12% of all available station records. A total of 23% of monthly records that had missing values could be estimated with high or moderate confidence. We sub-sampled our global database of more than 80,000 stations with various spatial filters, so that only the highest quality station for a given area was retained.</p> <p>Our contribution significantly enhances global data coverage compared to individual databases currently available. Even when accepting only the stations within the top two quality ranks in our combined database, and applying the coarsest spatial filter of one station per approximately 1,600 km², the remaining station count of more than 20,000 stations exceeds the largest alternative database (without a spatial filter applied) by more than 50%.</p> <p>The database contains a "Station Statistics" file with various flags indicating station quality and completeness of records. Monthly precipitation data is provided as one large file, but also broken down into regional files with less than one million rows each. Climate normal estimates for the 1961-1990 period, useful as a baseline prior to significant anthropogenic warming, are provided in multiple files with global coverage, but with different spatial filters applied that select the highest quality stations based for a global grid at different resolutions.</p>
Data from: sFDvent: a global trait database for deep-sea hydrothermal-vent fauna
Traits are increasingly being used to quantify global biodiversity patterns, with trait databases growing in size and number, across diverse taxa. Despite growing interest in a trait-based approach to the biodiversity of the deep sea, where the impacts of human activities (including seabed mining) accelerate, there is no single repository for species traits for deep-sea chemosynthesis-based ecosystems, including hydrothermal vents. Using an international, collaborative approach, we have compiled the first global-scale trait database for deep-sea hydrothermal-vent fauna - sFDvent (sDiv-funded trait database for the Functional Diversity of vents). We formed a funded working group to select traits appropriate to: i) capture the performance of vent species and their influence on ecosystem processes, and ii) compare trait-based diversity in different ecosystems. Forty contributors, representing expertise across most known hydrothermal-vent systems and taxa, scored species traits using online collaborative tools and shared workspaces. Here, we typify the sFDvent database, describe our approach, and evaluate its scope. Finally, we compare the sFDvent database to similar databases from shallow-marine and terrestrial ecosystems to highlight how the sFDvent database can inform cross-ecosystem comparisons. We also make the sFDvent database publicly available online by assigning a persistent, unique doi. 646 vent species names, associated location information (33 regions), and scores for 13 traits (in categories: community structure, generalist/specialist, geographic distribution, habitat use, life history, mobility, species associations, symbiont, and trophic structure). Contributor IDs, certainty scores, and references are also provided. Global coverage (grain size: ocean basin), spanning eight ocean basins, including vents on 12 mid-ocean ridges and 6 back-arc spreading centres. sFDvent includes information on deep-sea vent species, and associated taxonomic updates, since they were first discovered in 1977. Time is not recorded. The database will be updated every five years. Deep-sea hydrothermal-vent fauna with species-level identification present or in progress. .csv and MS Excel (.xlsx)
A vigiPoint characterisation of female versus male reports in VigiBase, the WHO global database of individual case safety reports
<p><strong>General information</strong></p> <p>This data is supplementary material to the paper by Watson <em>et al</em>. on sex differences in global reporting of adverse drug reactions [1]. Readers are referred to this paper for a detailed description of the context in which the data was generated. Anyone intending to use this data for any purpose should read the publicly available information on the VigiBase source data [2, 3]. The conditions specified in the caveat document [3] must be adhered to.</p> <p><b>Source dataset</b></p> <p>The dataset published here is based on analyses performed in VigiBase, the WHO global database of individual case safety reports [4]. All reports entered into VigiBase from its inception in 1967 up to 2 January 2018 with patient sex coded as either female or male have been included, except suspected duplicate reports [5]. In total, the source dataset contained 9,056,566 female and 6,012,804 male reports.</p> <p><strong>Statistical analysis</strong></p> <p>The characteristics of the female reports were compared to those of the male reports using a method called vigiPoint [6]. This is a method for comparing two or more sets of reports (here female and male reports) on a large set of reporting variables, and highlight any feature in which the sets are different in a statistically and clinically relevant manner. For example, patient age group is a reporting variable, and the different age groups 0 - 27 days, 28 days - 23 months et cetera are features within this variable. The statistical analysis is based on shrinkage log odds ratios computed as a comparison between the two sets of reports for each feature, including all reports without missing information for the variable under consideration. The specific output from vigiPoint is defined precisely below. Here, the results for 18 different variables with a total of 44,486 features are presented. 74 of these features were highlighted as so called vigiPoint key features, suggesting a statistically and clinically significant difference between female and male reports in VigiBase.</p> <p><strong>Description of published dataset</strong></p> <p>The dataset is provided in the form of a MS Excel spreadsheet (.xlsx file) with nine columns and 44,486 rows (excluding the header), each corresponding to a specific feature. Below follows a detailed description of the data included in the different columns.</p> <p><em>Variable</em>: This column indicates the reporting variable to which the specific feature belongs. Six of these variables are described in the original publication by Watson <em>et al</em>.: <em>country of origin</em>, <em>geographical region of origin</em>, <em>type of reporter</em>, <em>patient age group</em>, <em>MedDRA SOC</em>, <em>ATC level 2 of reported drugs</em>, <em>seriousness</em>, and <em>fatality </em>[1]. The remaining 12 are described here:</p> <ul> <li> <em>MedDRA HLGT</em> (high-level group term), <em>MedDRA HLT </em>(high-level term) and <em>MedDRA PT</em> (preferred term) are defined analogously to the <em>MedDRA SOC</em> (system organ class) [1], only at lower levels of the <a href="https://www.meddra.org/">MedDRA </a>(Medical Dictionary for Regulatory Activities) hierarchy. Here, MedDRA version 20.1 has been used.</li> <li> <em>ATC level 3 of reported drugs</em> is defined analogously to the variable <em>ATC level 2 of reported drugs </em>[1], only one step further down in the <a href="https://www.whocc.no/atc/structure_and_principles">ATC</a> (Anatomical Therapeutical Classification) hierarchy.</li> <li>The vigiGrade <i>completeness score</i> is a measure of how complete each report is with respect to certain report fields useful for causality assessment [7]. The completeness score has been dichotomised into two features, 'Above or equal to 0.8' and 'Below 0.8'. The maximum possible score for an individual report is 1.0.</li> <li>The <em>date of VigiBase entry </em>is simply the time when a report was entered into VigiBase. This variable is divided into 14 features that are either individual years or ranges of years.</li> <li> The <em>number of reported drugs</em> is the number of unique drugs that are coded on a report as either suspected, interacting, or concomitant. A drug is here defined as an entry at the preferred base (i.e. substance) level of the WHODRUG terminology. The variable is divided into four features: 'One drug', 'Two drugs', '3-5 drugs', and 'More than 5 drugs'.</li> <li> The <em>number of reported MedDRA PTs</em> is the number of unique MedDRA preferred terms that are coded as events on a report. This variable is divided into four features in exactly the same way as the reported drugs.</li> <li> A <i>reported drug</i> is a drug coded on a report as either suspected, interacting, or concomitant. As above, a drug is defined as an entry at the preferred base (i.e. substance) level of the WHODRUG terminology. This variable has almost 23,000 features, one for each drug that occurs in at least one female or one male report.</li> <li> The <em>type of report</em> indicates the type of individual case report. The vast majority belongs to the feature 'Spontaneous', but there are four other possible features for this variable.</li> </ul> <p>The <em>Variable </em>column can be useful for filtering the data, for example if one is interested in one or a few specific variables.</p> <p><i>Feature: </i>This column contains each of the 44,486 included features. The vast majority should be self-explanatory, or else they have been explained above, or in the original paper [1].</p> <p><em>Female reports</em> and <em>Male reports</em>: These columns show the number of female and male reports, respectively, for which the specific feature is present.</p> <p><em>Proportion among female reports</em> and <em>Proportion among male reports</em>: These columns show the proportions within the female and male reports, respectively, for which the specific feature is present. Comparing these crude proportions is the simplest and most intuitive way to contrast the female and male reports, and a useful complement to the specific vigiPoint output.</p> <p><em>Odds ratio</em>: The odds ratio is a basic measure of association between the classification of reports into female and male reports and a given reporting feature, and hence can be used to compare female and male reports with respect to this feature. It is formally defined as <em>a / (bc / d)</em>, where</p> <ul> <li> <em>a</em> is the number of female reports with the feature</li> <li> <em>b</em> is the number of female reports without the feature (excluding reports where the variable is missing)</li> <li> <em>c</em> is the number of male reports with the feature</li> <li> <i>d</i> is the number of male reports without the feature (excluding reports where the variable is missing).</li> </ul> <p>This crude odds ratio can also be computed as <em>(p<sub>female</sub> / (1-p<sub>female</sub>)) / (p<sub>male</sub> / (1-p<sub>male</sub>))</em>, where <em>p<sub>female</sub></em> and <em>p<sub>male</sub></em> are the proportions described earlier. If the odds ratio is above 1, the feature is more common among the female than the male reports; if below 1, the feature is less common among the female than the male reports. Note that the odds ratio can be mathematically undefined, in which case it is missing in the published data.</p> <p><em>vigiPoint score</em>: This score is defined based on an odds ratio with added statistical shrinkage, defined as (<em>a + k) / ((bc / d) + k)</em>, where <em>k</em> is 1% of the total number of female reports, or about 9,000. While the shrinkage adds robustness to the measure of association, it makes interpretation more difficult, which is why the crude proportions and unshrunk odds ratios are also presented. Further, 99% credibility intervals are computed for the shrinkage odds ratios, and these intervals are transformed onto a log<sub>2</sub> scale [6]. The vigiPoint score is then defined as the lower endpoint of the interval, if that endpoint is above 0; as the higher endpoint of the interval, if that endpoint is below 0; and otherwise as 0. The vigiPoint score is useful for sorting the features from strongest positive to strongest negative associations, and/or to filter the features according to some user-defined criteria.</p> <p><em>vigiPoint key feature</em>: Features are classified as vigiPoint key features if their vigiPoint score is either above 0.5 or below -0.5. The specific thereshold of 0.5 is arbitrary, but chosen to identify features where the two sets of reports (here female and male reports) differ in a clinically significant way.</p> <p><strong>References</strong></p> <ol> <li>Watson S, Caster O, Rochon PA, den Ruijter H. Reported adverse drug reactions in women and men: Aggregated evidence from globally collected individual case reports during half a decade. <em>EClinicalMedicine</em> 2019.</li> <li>Uppsala Monitoring Centre. <a href="https://www.who-umc.org/media/164772/guidelineusingvigibaseinstudies.pdf">Guideline for using VigiBase data in studies</a>.</li> <li>Uppsala Monitoring Centre. <a href="https://www.who-umc.org/media/164610/umc_caveat.pdf">Caveat document: Statement of reservations, limitations, and conditions relating to data released from VigiBase, the WHO global database of individual case safety reports (ICSRs)</a>.</li> <li>Lindquist M. VigiBase, the WHO Global ICSR Database System: Basic Facts. <i>The Drug Information Journal</i> 2008; <b>42</b>(5): 409-19.</li> <li>Norén GN, Orre R, Bate A, Edwards IR. Duplicate detection in adverse drug reaction surveillance. <em>Data Mining and Knowledge Discovery</em> 2007; <strong>14</strong>(3): 305-28.</li> <li>Juhlin K, Star K, Norén GN. A method for data-driven exploration to pinpoint key features in medical data and facilitate expert review. <i>Pharmacoepidemiology and Drug Safety</i> 2017; <b>26</b>(10): 1256-65.</li> <li>Bergvall T, Norén GN, Lindquist M. vigiGrade: A tool to identify well-documented individual case reports and highlight systematic data quality issues. <em>Drug Safety </em>2014; <strong>37</strong>(1): 65-77.</li> </ol>
Supplementary material 9 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Unique continent and island hit combinations in Insecta dataset
Supplementary material 8 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Number of distinct continent or island groupings recovered per Insecta BIN
Supplementary material 7 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
List of intercontinental and island Insecta BINs with identification metadata
Supplementary material 6 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Intercontinental and island records for targeted Platygastroidea in GBIF and the literature
Supplementary material 4 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Platygastroidea COI dataset NJ tree.tre
Supplementary material 2 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Randomized BOLD BINs for validation of the Insecta dataset
Supplementary material 19 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Randomized BOLD BINs for validation of the Araneae dataset
Supplementary material 18 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
List of intercontinental and island Araneae BINs with identification metadata
Supplementary material 3 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Intercontinental and island Platygastroidea COI dataset alignment
Supplementary material 17 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Platygastroidea BIN identifications using digital morphology infrastructure
Supplementary material 16 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Pairwise geographic hit comparisons for the Trissolcus BIN, GBIF, and literature dataset
Supplementary material 14 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Pairwise geographic hit comparisons for the Synopeas BIN, GBIF, and literature dataset
Supplementary material 13 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Pairwise geographic hit comparisons for the Platygaster BIN, GBIF, and literature dataset
Supplementary material 12 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Pairwise geographic hit comparisons for the Platygastroidea BIN, GBIF, and literature dataset
Supplementary material 15 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Pairwise geographic hit comparisons for the Telenomus BIN, GBIF, and literature dataset
Supplementary material 11 from: Moore MR, Talamas EJ, Bremer JS, McGathey N, Fulton JC, Lahey Z, Awad J, Roberts CG, Combee LA (2023) Mining biodiversity databases establishes a global baseline of cosmopolitan Insecta mOTUs: a case study on Platygastroidea (Hymenoptera) with consequences for biological control programs. NeoBiota 88: 169-210. https://doi.org/10.3897/neobiota.88.106326
Summary of intercontinental and island Platygastroidea BINs
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.