Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,505
datasets available to search
ShareScore release 0.9.0
Dataset results
3,505 results for “completeness”
The local extinction of Cedrus atlantica in the Iberian Peninsula could have been completed due to biological interaction
<p>This data set is used to explore the possibility that <em>Cedrus atlantica</em> (Endl.) Carrière and <em>Pinus nigra</em> Arnold could have interacted in the past, mutually excluding each other in the areas with suitable conditions for both species and, where, ultimately, the one that was most competitive would remain. The species show very well differenciated niches and a distribution of their habitats segregated by continents (<em>P. nigra</em> in Europe and <em>C. atlantica</em> in Africa), which responds to differences in climatic affinities. However, the contact of their distributions in bordering areas suggests that <em>C. atlantica</em> maintained its presence in the Iberian Peninsula until recent times, and that <em>P. nigra</em> could have displaced it due to its higher prevalence on the continent.</p>
Supplementary dataset to publication: Complete Genome Sequence of Ovine Mycobacterium avium subsp. paratuberculosis Strain JIII-386 (MAP-S/type III) and Its Comparison to MAP-S/type I, MAP-C, and M. avium Complex Genomes.
<p>This is the modified supplemented material to the publication “Complete genome sequence of ovine Mycobacterium avium subsp. paratuberculosis strain JIII-386 (MAP-S/type III) and its comparison to MAP-S/type I, MAP-C, and M. avium complex genomes”.</p> <p>The complete circular genome of Mycobacterium avium subsp. paratuberculosis (MAP) strain JIII-386 from Germany, closed by Nanopore technology in this study, was presented and compared with the draft genome of JIII-386, previously published in [doi:10.1093/gbe/ew154], the closed genome of the MAP-S/type I strain Telford, the MAP-S/type III draft genome of strain S397, twelve closed MAP-C (type II) strains and eight closed Mycobacterium avium (M. a.) strains of subsp. hominissuis (MAH) and subsp. avium (MAA). Structural comparisons clearly revealed the mosaic nature of MAP genomes, the differences between MAP subtypes I, II and III, and the higher diversity of MAP-S compared to MAP-C genomes. </p> <p>The material provides a wealth of detailed results from these analyses and comparisons. These include a list of identified ncRNA and Riboswitches, as well as additional genes in finished JIII-386, the gene content of identified prophage regions, copy number of identified transposable elements and a list of selected virulence-associated genes in the different MAP-type (I - III) strains. The genomic islands identified and included genes along with their predicted functions were presented for six MAP genomes (belonging to MAP-S/type I and III, and MAP-C), one MAH genome and one MAA genome. One table shows the corresponding genomic islands in the genomes of JIII-386, Telford and three MAP-C genomes. Furthermore, homologous genes of known MAP-S specific Large Sequence Polymorphisms regions (LSP<sup>S</sup> = LSP-S) were recorded in different MAP-S type strains, one MAH and one MAA strain, as well as genes of deletions #1 (LSP<sup>A</sup>-20), #2, and s-delta-1, previously described as MAP-S-specific deletions, their presence or absence in 3 MAP-S, 12 MAP-C, 4 MAH, and 4 MAA strains were listed. Different presence or absence of genes, but also identified frameshifts or disruptions of various virulence-associated genes could lead to the different MAP-type specific phenotypic characteristics. Comprehensive core and pan genome analyses (results listed in six tables) revealed unique genes and genes likely to have been acquired by horizontal gene transfer in different MAP types and subtypes, but also emphasized the highly conserved and close relationship, and the complex evolution of M. a. strains.</p> <p> </p>
Dataset for the IntoValue 1 + 2 studies on results dissemination from clinical trials conducted at German university medical centers completed between 2009 and 2017
<p>The IntoValue dataset contains clinical trials conducted at one of 35 German UMCs and registered on ClinicalTrials.gov or the German Clinical Trials Registry (DRKS). All trials were reported as complete between 2009 and 2017 on the trial registry at the time of data collection. The dataset also includes a results publication found via manual searches; if multiple results publications were found, the earliest was included.</p> <p>Trials were associated with a German UMC by searching for trials with a UMC listed as responsible party or lead sponsor, or with a principle investigator (PI) from a UMC ('lead_city'). Version 1 additionally includes trials with a UMC only as a facility (`facility_city`). A lookup table of regular expressions used to identify German UMCs is available at <a href="https://github.com/quest-bih/IntoValue2/blob/master/data/1_sample_generation/city_search_terms.csv">https://github.com/quest-bih/IntoValue2/blob/master/data/1_sample_generation/city_search_terms.csv</a>.</p> <p>Trials include all interventional studies and are not limited to investigational medical product trials, as regulated by the EU's Clinical Trials Directive or Germany's Arzneimittelgesetz (AMG) or Novelle des Medizinproduktegesetzes (MPG).</p> <p>DRKS data were searched (pre-filtered for completion years and study status as well as Germany as 'Country of recruitment') and downloaded as CSVs from the DRKS website (<a href="https://www.drks.de/">https://www.drks.de/</a>). ClinicalTrials.gov data were downloaded downloaded as pipe files from Clinical Trials Transformation Initiative (CTTI) Aggregate Content of ClinicalTrials.gov (AACT) (<a href="https://aact.ctti-clinicaltrials.org/pipe_files">https://aact.ctti-clinicaltrials.org/pipe_files</a>). DRKS and ClinicalTrials.gov use different terminology for various trial aspects, such as phase and masking; these different levels are captured in the data dictionary as `levels_drks` and `levels_ctgov`. For later analyses requiring parity across registries, levels for some variables were collapsed and a lookup table is provided in `iv_data_lookup_registries.csv`.</p> <p>These data were generated and used for two publications (Wieschowski et al., 2019; Riedel et al. 2021) and therefore comprises two versions (indicated as `iv_version`).</p> <p>For version 1, registry data was collected on April 17, 2017 from ClinicalTrials.gov and on July 27, 2017 for DRKS and was limited to trials with a completion date on DRKS and primary completion date on ClinicalTrials.gov between 2009 and 2013. Version 1 manual searches for results publications were conducted from 2017-07-01 to 2017-12-01.<br> For version 2, registry data was collected on June 3, 2020 and was limited to trials with a completion date on DRKS and ClinicalTrials.gov between 2014 and 2017. Version 2 manual searches for results publications were conducted from 2020-07-01 to 2020-09-01.</p> <p>Raw registry data for versions 1 and 2 is available in `raw-registries.zip`.</p> <p>Publication identifiers (DOI, PMID, URL) were manually entered during the publication search and then further enhanced using the API of Internet Archive's open-source Fatcat catalog of research publications, to add PMIDs based on DOIs, and vice versa.</p> <p>Manual search steps differed slightly in the two versions and are indicated and described in `identification_step`.<br> Version 1 includes trials with a German UMC as either a `lead_city` or a `facility_city`, whereas version 2 is limited to trials a German UMC as a `lead_city`.</p> <p>Each row indicates a single trial registration. Due to changes in completion dates, some trials are duplicated between versions as indicated in `is_dupe`. Cross-registered trials were manually deduplicated, and some cross-registered duplicates remain (e.g., DRKS00004156 and NCT00215683) and are not indicated in the dataset.</p> <p>All dates are provided as `yyyy-mm-dd`.</p> <p>Additional documentation on each variable (type, description, levels) is provided in `iv_data_dictionary.csv`.</p> <p>Additional information on the project and methods for generating the dataset is available in associated publications and at the project's OSF page (<a href="https://osf.io/98j7u/">https://osf.io/98j7u/</a>). Code for the project is available at <a href="https://github.com/quest-bih/IntoValue2">https://github.com/quest-bih/IntoValue2</a>.</p> <p><strong>References:</strong></p> <p>Wieschowski, S., Riedel, N., Wollmann, K., Kahrass, H., Müller-Ohlraun, S., Schürmann, C., Kelley, S., Kszuk, U., Siegerink, B., Dirnagl, U., Meerpohl, J., & Strech, D. (2019). Result dissemination from clinical trials conducted at German university medical centers was delayed and incomplete. Journal of Clinical Epidemiology, 115, 37–45. <a href="https://doi.org/10.1016/j.jclinepi.2019.06.002">https://doi.org/10.1016/j.jclinepi.2019.06.002</a></p> <p>Riedel, N., Wieschowski, S., Bruckner, T., Holst, M. R., Kahrass, H., Nury, E., Meerpohl, J. J., Salholz-Hillel, M., & Strech, D. (2021). Results dissemination from completed clinical trials conducted at German university medical centers remained delayed and incomplete. The 2014-2017 cohort. Journal of Clinical Epidemiology, 0(0). <a href="http://doi.org/10.1016/j.jclinepi.2021.12.012">https://doi.org/10.1016/j.jclinepi.2021.12.012</a><br> </p>
Nearshore high-frequency temporal water quality observations and process-based modeling of aquatic ecosystem metabolism in Lake Tahoe completed by members of the Blaszczak Lab at the University of Nevada Reno, 2021-2023
The overarching goal of this project was to develop a process-based understanding of how watershed-to-lake connections drive nearshore productivity dynamics in a large oligotrophic mountain lake (Lake Tahoe). We addressed this goal through a combined approach of high-frequency sensor deployment and maintenance, ecosystem metabolism modeling, laboratory incubations, and routine monitoring of water chemistry and other parameters. The data we collected as part of this project and the ecosystem metabolism estimates we generated demonstrate how variable ecosystem productivity is in time and space in the nearshore of Lake Tahoe. Although maintenance of the sensor arrays during the exceptional winter of 2023 was challenging, we were able to capture the data necessary to estimate a complete time series of metabolic activity across two years with very different hydroclimatic conditions. Throughout this project we accomplished the following: 1. We generated over two years of daily estimates of ecosystem metabolism (gross primary productivity, ecosystem respiration, and net ecosystem productivity) from multiple locations on both the east and west shores of the lake and from areas in close proximity to and far away from stream water inflows. 2. We measured ammonium (NH4+) and nitrate (NO3-) concentrations in surface water samples from both Glenbrook and Blackwood creeks and the nearshore of Lake Tahoe for over two years. 3. We quantified rates of NH4+ and NO3- uptake in benthic samples of the dominant substrate type collected during peak streamflow, the receding limb, and baseflow conditions in 2023 from multiple locations in the nearshore using established laboratory incubation methods. 4. Finally, we used a combination of time series models and structural equation modeling to integrate our results and improve understanding of the direct and indirect effects of hydroclimatic variability on observed patterns in ecosystem metabolism in the nearshore. See this git code repository
Darwin: an amino acid sequence collection of complete proteomes from eukaryotes with different phylogenetic affinities (v. 03_2020_137)
<p><strong>Background</strong></p> <p>Every time we find an interesting gene in an organism of interest, the first question is often “how widely is this gene distributed in the eukaryotic kingdom?”. Naturally, one could use NCBI BLAST search against the non-redundant sequence database provided by GenBank to answer this question. However, it can be cumbersome to parse the results and assign them to taxonomic units. It is also not straightforward to get an overview of which eukaryotic groups are represented in the results. Top BLAST hits can be crowded with sequences from closely-related organisms making it difficult gain an overview of the overall distribution across eukaryotes. To streamline this process, we developed an in-house database of complete eukaryotic proteomes. We tagged each sequence with a eukaryotic group handle (two-character symbol) and combined them into a single data set searchable by standalone BLAST on one’s own computer. We named this data set “Darwin” to reflect the diverse nature of the sequences it contains. </p> <p><strong>Methods</strong></p> <p>We downloaded predicted proteomes in FASTA format from different sources such as GenBank, Joint Genome Institute (Depart of Energy, USA), Broad Institute (Massachusetts Institute of Technology, USA), Phytozome and a number of other specialized websites catering for a specific organism such as the Arabidopsis Information Resource (TAIR), or the Saccharomyces Genome Database (SGD). All the organisms we included in Darwin are listed in Table 1. To reduce redundancy, we took care not to include the same species more than once unless subspecies were known to show wide diversity. Each sequence header was tagged with a eukaryotic group handle composed of two-character symbols (based on Keeling <em>et al</em>., 2005). These handles clearly appear in BLAST output and can be parsed easily. We combined sequences from all proteomes into a single data set and named it “Darwin”.</p> <p><strong>Results</strong></p> <p>The current version of Darwin (v. 03_2020_137) contains 2,601,132 amino acid sequences from 137 eukaryotes (Table 1, Data file 1). The sizes of the proteomes were diverse, ranging from ~4000 sequences in some alveolates to 60,000-76,000 in plants. Darwin represents most of the supergroups of eukaryotic kingdom described in Keeling <em>et al.,</em> (2005) except those in Rhizaria whose genomes were not available at the time of data set construction. The data set contains larger numbers of proteomes from fungi and plants reflecting areas of interest in our group. </p> <p><strong>Conclusions</strong></p> <p>Darwin is provided as a text fasta file that can be formatted for BLAST searches on standalone computers. The results from the BLAST searches can be parsed to determine how widely a gene of interest is distributed among different eukaryotes. Simple counting of the eukaryotic group handles would also yield an overview of the distribution across taxa. Darwin is also useful for rapidly finding out whether a gene is missing in particular taxa.</p> <p><strong>Reference</strong></p> <p>Keeling PJ, Burger G, Durnford DG, Lang BF, Lee RW, Pearlman RE, Roger AJ, Gray MW (2005) The tree of eukaryotes. <em>Trends Ecol. Evol.</em> <strong>20:</strong> 670-676</p>
Global Human Settlement Layer per zoom-level 18 Quadtree tile for selected countries as Spatialite database with OpenStreetMap building completeness assessment
<p>This Spatialite database contains the built-up area of the Global Human Settlement Layer (GHSL) per zoom-level 18 Quadtree tile. Additionally, it provides a comparison of the GHSL with buildings in OpenStreetMap: For each tile the built-up ratio between the building footprints and the GHSL is given and a binary completeness assessment (buildings complete, not complete) is provided for easy use. This dataset was created using the obmgapanalysis tool: https://git.gfz-potsdam.de/dynamicexposure/openbuildingmap/obmgapanalysis</p>
WikiLinkGraphs: A complete, longitudinal and multilanguage dataset of the Wikipedia link networks
<p>This dataset contains yearly snapshots of the Wikipedia's internal link network for the 9 largest language edition (de, en, es, fr, it, nl, pl, ru, sv). The dataset spans over 17 years, from the creation of Wikipedia in 2001 to March 2018. The snapshots are taken on March 1st of every year.</p> <p>The graphs include the links extract from the wikitext of each page (i.e in the form [[wikilink]]). Links transcluded from templates are not included. Redirects are resolved to their target page.</p> <p>More detailed information and supporting datasets are available at: http://disi.unitn.it/~consonni/datasets/.</p> <p><strong>IMPORTANT NOTICE</strong></p> <p>Gzipped files are compressed two times by Zenodo, the MD5 provided by Zenodo and the SHA512 sums provided in the `.sha512sums.txt` files, match with the files compressed once. In other words, when you download a `.gz` file save it as `.gz.gz`, uncompress it once and it should match both the MD5 provided by Zenodo and the SHA512 sum provided by us. We have opened a bug report for this behavior on Zenodo's repository at: https://github.com/zenodo/zenodo/issues/1705</p> <p> </p>
Dataset for the publication "Implementation of an exact completing method of generation for face-milled spiral bevel gears with uniform depth taper"
<p>This dataset contains geometric and graphics data associated with the referenced paper, enabling the reproduction of the conducted research. </p>
Complete Rxivist dataset of scraped biology preprint data
<p><a href="https://rxivist.org">rxivist.org</a> allowed readers to sort and filter the tens of thousands of preprints posted to <a href="https://www.biorxiv.org">bioRxiv</a> and <a href="https://www.medrxiv.org">medRxiv</a>. Rxivist used a custom web crawler to index all papers posted to those two websites; this is a snapshot of Rxivist the production database. The version number indicates the date on which the snapshot was taken. See the included "README.md" file for instructions on how to use the "rxivist.backup" file to import data into a PostgreSQL database server.</p> <p>Please note this is a different repository than the one used for <a href="https://www.biorxiv.org/content/early/2019/01/13/515643">the Rxivist manuscript</a>—that is in <a href="https://doi.org/10.5281/zenodo.2465689">a separate Zenodo repository</a>. You're welcome (and encouraged!) to use this data in your research, but <strong>please cite our paper, now published <a href="https://doi.org/10.7554/eLife.45133">in <em>eLife</em></a>.</strong></p> <p>Previous versions are also available pre-loaded into Docker images, available at <a href="https://hub.docker.com/r/blekhmanlab/rxivist_data">blekhmanlab/rxivist_data</a>.</p> <p><strong>Version notes:</strong></p> <ul> <li><strong>2023-03-01</strong> <ul> <li>The final Rxivist data upload, more than four years after the first and encompassing 223,541 preprints posted to bioRxiv and medRxiv through the end of February 2023.</li> </ul> </li> <li><em><strong>2020-12-07***</strong></em> <ul> <li>In addition to bioRxiv preprints, <em><strong>the database now includes all medRxiv preprints as well</strong></em>. <ul> <li>The website where a preprint was posted is now recorded in a <strong>new field</strong> in the "articles" table, called "<strong>repo</strong>".</li> </ul> </li> <li>We've significantly refactored the web crawler to take advantage of developments with the bioRxiv API. <ul> <li>The main difference is that preprints flagged as "published" by bioRxiv are no longer recorded on the same schedule that download metrics are updated: The Rxivist database should now record published DOI entries the same day bioRxiv detects them.</li> </ul> </li> <li>Twitter metrics have returned, for the most part. Improvements with the Crossref Event Data API mean we can once again tally daily Twitter counts for all bioRxiv DOIs. <ul> <li>The "crossref_daily" table remains where these are recorded, and daily numbers are now up to date.</li> <li>Historical daily counts have also been re-crawled to fill in the empty space that started in October 2019.</li> <li>There are still several gaps that are more than a week long due to missing data from Crossref.</li> <li>We have recorded available Crossref Twitter data for all papers with DOI numbers starting with "10.1101," which includes all medRxiv preprints. However, <strong>there appears to be almost no Twitter data available for medRxiv preprints</strong>.</li> </ul> </li> <li>The download metrics for article id 72514 (DOI 10.1101/2020.01.30.927871) were found to be out of date for February 2020 and are now correct. This is notable because article 72514 is the most downloaded preprint of all time; we're still looking into why this wasn't updated after the month ended.</li> </ul> </li> <li><strong>2020-11-18</strong> <ul> <li>Publication checks should be back on schedule.</li> </ul> </li> <li><strong>2020-10-26</strong> <ul> <li>This snapshot fixes most of the data issues found in the previous version. Indexed papers are now up to date, and download metrics are back on schedule. <em>The check for publication status remains behind schedule</em>, however, and the database may not include published DOIs for papers that have been flagged on bioRxiv as "published" over the last two months. Another snapshot will be posted in the next few weeks with updated publication information.</li> </ul> </li> <li><strong>2020-09-15</strong> <ul> <li>A crawler error caused this snapshot to exclude all papers posted after about August 29, with some papers having download metrics that were more out of date than usual. The "last_crawled" field is accurate.</li> </ul> </li> <li><strong>2020-09-08</strong> <ul> <li>This snapshot is misconfigured and will not work without modification; it has been replaced with version 2020-09-15.</li> </ul> </li> <li><strong>2019-12-27</strong> <ul> <li>Several dozen papers did not have dates associated with them; that has been fixed.</li> <li>Some authors have had two entries in the "authors" table for portions of 2019, one profile that was linked to their ORCID and one that was not, occasionally with almost identical "name" strings. This happened after bioRxiv began changing author names to reflect the names in the PDFs, rather than the ones manually entered into their system. These database records are mostly consolidated now, but some may remain.</li> </ul> </li> <li><strong>2019-11-29</strong> <ul> <li>The Crossref Event Data API remains down; Twitter data is unavailable for dates after early October.</li> </ul> </li> <li><strong>2019-10-31</strong> <ul> <li>The Crossref Event Data API is still <a href="https://status.crossref.org/">experiencing problems</a>; the Twitter data for October is incomplete in this snapshot.</li> <li>The README file has been modified to reflect changes in the process for creating your own DB snapshots if using the newly released PostgreSQL 12.</li> </ul> </li> <li><strong>2019-10-01</strong> <ul> <li>The Crossref API is back online, and the "crossref_daily" table should now include up-to-date tweet information for July through September.</li> <li>About 40,000 authors were removed from the author table because the name had been removed from all preprints they had previously been associated with, likely because their name changed slightly on the bioRxiv website ("John Smith" to "J Smith" or "John M Smith"). The "author_emails" table was also modified to remove entries referring to the deleted authors. The web crawler is being updated to clean these orphaned entries more frequently.</li> </ul> </li> <li><strong>2019-08-30</strong> <ul> <li>The Crossref Event Data API, which provides the data used to populate the table of tweet counts, has not been fully functional since early July. While we are optimistic that accurate tweet counts will be available at some point, the sparse values currently in the "crossref_daily" table for July and August should not be considered reliable.</li> </ul> </li> <li><strong>2019-07-01</strong> <ul> <li>A new "institution" field has been <a href="https://github.com/blekhmanlab/rxivist/commit/1cb570703085841e80cc3073af445bc86f0cbb63#diff-a04b1c1a66d16f9b3acfed9b9d2128c5">added</a> to the "article_authors" table that stores each author's institutional affiliation <em>as listed on that paper</em>. The "authors" table still has each author's most recently observed institution. <ul> <li>We began collecting this data in the middle of May, but it has not been applied to older papers yet.</li> </ul> </li> </ul> </li> <li><strong>2019-05-11</strong> <ul> <li>The README was updated to correct a link to the Docker repository used for the pre-built images.</li> </ul> </li> <li><strong>2019-03-21</strong> <ul> <li>The license for this dataset has been changed to <a href="https://creativecommons.org/licenses/by/4.0/">CC-BY</a>, which allows use for any purpose and requires only attribution.</li> <li>A new table, "publication_dates," has been added and will be continually updated. This table will include an entry for each preprint that has been published externally for which we can determine a date of publication, based on data from Crossref. (This table was previously included in the "paper" schema but was not updated after early December 2018.)</li> <li>Foreign key constraints have been added to almost every table in the database. This should not impact any read behavior, but anyone writing to these tables will encounter constraints on existing fields that refer to other tables. Most frequently, this means the "article" field in a table will need to refer to an ID that actually exists in the "articles" table.</li> <li>The "author_translations" table has been removed. This was used to redirect incoming requests for outdated author profile pages and was likely not of any functional use to others.</li> <li>The "README.md" file has been renamed "1README.md" because Zenodo only displays a preview for the file that appears first in the list alphabetically.</li> <li>The "article_ranks" and "article_ranks_working" tables have been removed as well; they were unused.</li> </ul> </li> <li><strong>2019-02-13.1</strong> <ul> <li>After consultation with bioRxiv, the "fulltext" table will not be included in further snapshots until (and if) concerns about licensing and copyright can be resolved.</li> <li>The "docker-compose.yml" file was added, with corresponding instructions in the README to streamline deployment of a local copy of this database.</li> </ul> </li> <li><strong>2019-02-13</strong> <ul> <li>The redundant "paper" schema has been removed.</li> <li>BioRxiv has begun making the full text of preprints available online. Beginning with this version, a new table ("fulltext") is available that contains the text of preprints that have been processed already. <strong>The format in which this information is stored may change in the future</strong>; any digression will be noted here.</li> <li>This is the first version that has <a href="https://cloud.docker.com/u/blekhmanlab/repository/docker/blekhmanlab/rxivist_data">a corresponding Docker image</a>.</li> </ul> </li> </ul>
American Residential Macrosystems - Complete municipal ordinance documents across six U.S. cities, 2017-2019
These data files are the complete city codes, or municipal ordinances (n=156), across the metropolitan regions of Los Angeles, CA; Phoenix, AZ; Miami, FL; Baltimore, MD; Boston, MA; Minneapolis/St. Paul, MN. The documents were gathered for the specific purposes of a content analysis of how cities regulate residential landscapes; however, the documents include regulations on the books for the municipalities sampled for this project.
The Contributionsof Eye Gaze Fixations and Target-Lure Similarity to Behavioral and fMRI Indices of Pattern Separation and Pattern Completion
Open the record for dataset details and reuse information.
Marker files for assessing Prochlorococcus genome completness
<p>Use these files with CheckM to assess Prochlorococcus genome completeness:</p> <p><em>pro-marker-checkm-refined.txt</em></p> <p>List of custom PFAM and TIGRFAM identifiers used to assess genome completeness in <em>Prochlorococcus</em> using checkM</p> <p><em>pro-marker-checkm-refined.ms</em></p> <p>CheckM marker file of custom PFAM and TIGRFAM identifiers used to assess genome completeness in <em>Prochlorococcus</em>.</p> <p><em>pro-marker-checkm-refined.hmm</em></p> <p>HMM file of custom Prochlorococcus markers for use with CheckM and hmmer3</p>
Deliverable D-IA.2.2.OH-Harmony-Cap.2.1: Completed Pilot Survey
<p>This is a public deliverable of One Health EJP Joint Research Project,<strong><em> </em></strong><strong><em>Integrative Action-2.2,</em></strong> <strong>OH-HARMONY-CAP: </strong>One Health Harmonisation of Protocols for the Detection of Foodborne Pathogens and AMR Determinants. <a href="https://onehealthejp.eu/jip-oh-harmony-cap/"><strong>https://onehealthejp.eu/jip-oh-harmony-cap/</strong></a></p> <p>The purpose is to develop an integrated One Health map (OHLabCap ) of the levels of system capability/capacity/interoperability for each of the EU MS that is repeatable and sustainable. The first step is to develop of a pilot survey, targeting NRLs and the primary diagnostic services (primary sector). The pilot survey covers: 1. six priority bacteria and ten priority parasites have been chosen, as model organisms, together with the antimicrobial resistance (AMR) testing of <em>Salmonella</em> and <em>Campylobacter</em>. 2. 63 questions incorporating capability, capacity, and interoperability</p>
Dataset of 'Complete flow characterization from snapshot PIV, fast probes and physics-informed neural networks'
<p>Dataset of the article 'Complete flow characterization from snapshot PIV, fast probes and physics-informed neural networks' (https://doi.org/10.1016/j.cma.2023.116652). The codes processing data here are on https://github.com/AlvaroMS90/Complete-flow-characterization-from-snapshot-PIV-fast-probes-and-physics-informed-neural-networks.</p> <p>This project has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation program (grant agreement No 949085) and by MCIN/AEI /10.13039/501100011033 and the European Union ‘NextGenerationEU/PRTR’ as part of the grant FJC2020-044342-I.</p>
TokTrack: A Complete Token Provenance and Change Tracking Dataset for the English Wikipedia
<p><strong>Fixes in version 1.1 (= Zenodo's "version 2")</strong></p> <p>*In 20161101-revisions-part1-12-1728.csv, missing first data line is added.</p> <p>*In Current_content and Deleted_content files, some token values ('str' column) which contain regular quotes ('"') are fixed.</p> <p>*In Current_content and Deleted_content files, some wrong revision ID values for 'origin_rev_id', 'in' and 'out' columns are fixed.</p> <p> ------</p> <p><strong>This dataset contains every instance of all tokens (≈ words) ever written in undeleted, non-redirect English Wikipedia articles until October 2016, in total 13,545,349,787 instances. Each token is annotated with (i) the article revision it was originally created in, and (ii) lists with all the revisions in which the token was ever deleted and (potentially) re-added and re-deleted from its article, enabling a complete and straightforward tracking of its history.</strong></p> <p>This data would be exceedingly hard to create by an average potential user as it is (i) very expensive to compute and as (ii) accurately tracking the history of each token in revisioned documents is a non-trivial task. <br> Adapting a state-of-the-art algorithm, we have produced a dataset that allows for a range of analyses and metrics, already popular in research and going beyond, to be generated on complete-Wikipedia scale; ensuring quality and allowing researchers to forego expensive text-comparison computation, which so far has hindered scalable usage.</p> <p>This dataset, its creation process and use cases are described in a dedicated dataset paper of the same name, published at the ICWSM 2017 conference. In this paper, we show how this data enables, on token level, computation of provenance, measuring survival of content over time, very detailed conflict metrics, and fine-grained interactions of editors like partial reverts, re-additions and other metrics.</p> <p>Tokenization used: https://gist.github.com/faflo/3f5f30b1224c38b1836d63fa05d1ac94</p> <p>Toy example for how the token metadata is generated: <br> https://gist.github.com/faflo/8bd212e81e594676f8d002b175b79de8</p> <p><strong>Be sure to read the ReadMe.txt or - even more detailed - the supporting paper which is referenced under "related identifiers".</strong></p>
The complete reference genome for grapevine (Vitis vinifera L.) genetics and breeding
<div>PN40024, a highly homozygous inbred line originating from ‘Helfensteiner’, was used for T2T genome assembly. In total, 21 Gb (21 024 461 524 bp, ∼42× coverage) HiFi reads were generated by the PacBio platform. For the preliminary assembly, hifiasm was used to assemble the HiFi reads. We then used MUMmer and the 12X.v0 genome version (V. vinifera genome assembly 12X.v0 to order the 38 contigs into 19 chromosomes.</div> <p> The PN_T2T genome size was finally generated (494.87 Mb), being 69 Mb longer than 12X.v0 using the same statistical method. The k-mer metric was used to evaluate genomic homozygosity, estimated at 99.8%. The BUSCO for this genome is up to 98.5%.</p> <p>The PN40024.T2T genome assembly: PN.fa</p> <p>The PN40024.T2T gene annotation: PN_T2T.v5.1.gff3</p> <p>The PN40024.T2T TE annotation: PN_T2T_TE.gff</p> <p>The PN40024.T2T centromere annotation: PN.trf.gff3</p> <p>The PN40024.T2T protein sequence: PN_protein.fa</p> <p>The PN40024.T2T cds sequence: PN40024.cds.fa</p> <p>Comparison of gene annotation among PN_T2T and PN_T2T.v5.1, 12X.v0, 12X.v2, PN40024.v4, PN40024.v4.1: correlation.list.txt</p> <p>Mitochondrial assembly sequence of PN40024: PN_T2T_mit.fa</p> <p>Annotation of mitochondrial assembly for PN40024:PN_T2T_mit.gff3</p> <p>Chloroplast assembly sequence of PN40024: PN_T2T_chl.fa</p> <p>Annotation of chloroplast assembly for PN40024: PN_T2T_chl.gff3</p> <p>Citation: </p> <p>Please cite this paper when using the data of PN_T2T for your publications.</p> <p>Xiaoya Shi, Shuo Cao, Xu Wang, Siyang Huang, Yue Wang, Zhongjie Liu, Wenwen Liu, Xiangpeng Leng, Yanling Peng, Nan Wang, Yiwen Wang, Zhiyao Ma, Xiaodong Xu, Fan Zhang, Hui Xue, Haixia Zhong, Yi Wang, Kekun Zhang, Amandine Velt, Komlan Avia, Daniela Holtgräwe, Jérôme Grimplet, José Tomás Matus, Doreen Ware, Xinyu Wu, Haibo Wang, Chonghuai Liu, Yuling Fang, Camille Rustenholz, Zongming Cheng, Hua Xiao, Yongfeng Zhou, The complete reference genome for grapevine (<em>Vitis vinifera</em> L.) genetics and breeding, <em>Horticulture Research</em>, Volume 10, Issue 5, May 2023, uhad061, <a href="https://doi.org/10.1093/hr/uhad061">https://doi.org/10.1093/hr/uhad061</a></p>
Learning to Give a Complete Argument with a Conversational Agent: An Experimental Study in Two Domains of Argumentation
<p>This data is collected to find out how having a conversation with our agent affects argumentation. To model the arguments, we used Toulmin's model of argument. Based on the model, a good argument contains 3 different parts: 1. Claim, 2. Warrant, 3. Evidence. Based on Toulmin's model, these three components are the core components of arguments. This dataset has been collected during a between-subject experiment in which the treatment groups first talked to our agent in Task 1 and received feedback on faulty structural arguments and then did Task 2 and 3 which were answering a question on the same and completely different topic in comparison to Task 1.</p>
Complete datasets and code for "Hungry or angry? Experimental evidence for the effects of food availability on two measures of stress in developing wild raptor nestlings"
<p><strong>Abstract</strong></p> <p>Food shortage challenges the development of nestlings; yet, to cope with this stressor, nestlings can induce stress responses to adjust metabolism or behaviour. Food shortage also enhances the antagonism between siblings, but it remains unclear whether the stress response induced by food shortage operates via the individual nutritional state or via the social environment experienced. In addition, the understanding of these processes is hindered by the fact that effects of food availability often co-vary with other environmental factors. We used a food supplementation experiment to test the effect of food availability on two complementary stress measures, feather corticosterone (CORTf) and Heterophil/Lymphocyte-ratio (H/L) in developing red kite (Milvus milvus) nestlings, a species with competitive brood hierarchy. By statistically controlling for the effect of food supplementation on the nestlings’ body condition, we disentangled the effects of food and ambient temperature on nestlings during development. Experimental food supplementation increased body condition, and both CORTf and H/L were reduced in nestlings of high body condition. Additionally, CORTf decreased with age in non-supplemented nestlings. H/L decreased with age in all nestlings and was lower in supplemented last-hatched nestlings compared to non-supplemented ones. Ambient temperature showed a negative effect on H/L. Our results indicate that food shortage increases the nestlings’ stress levels through both, a reduced food intake affecting nutritional state and the nestlings’ social environment. Thus, food availability in conjunction with ambient temperature shape between- and within nest differences in stress load, which may have carry-over effects on behaviour and performance in further life-history stages.</p>
A complete map of specificity encoding for a partially fuzzy protein interaction
<p>All data required to run analyses for "A complete map of specificity encoding for a partially fuzzy protein interaction". Please see <a href="https://github.com/lehner-lab/fuzzy_specificity">https://github.com/lehner-lab/fuzzy_specificity</a> for instructions. </p>
DEM Intercomparison eXercise (DEMIX) - Maps of completeness criteria scores for global DEMs
<h2>Introduction</h2> <p>This introduction gives a brief overview of the context in which the dataset has been produced. Readers curious about the detailed standards and procedures described in this section are encouraged to open the resources linked to this dataset.</p> <h3>The Digital Elevation Model Intercomparison eXercise (DEMIX)</h3> <p>This work is part of the Digital Elevation Model Intercomparison eXercise (DEMIX), initiated by the <a href="https://ceos.org/ourwork/workinggroups/wgcv/current-activites/#:~:text=DEMIX%3A%20Digital%20Elevation%20Model%20Intercomparison,elevation%20model%20for%20their%20application.">Committee on Earth Observation Satellites (CEOS)</a>. This initiative aims at "<a href="https://isprs-archives.copernicus.org/articles/XLIII-B4-2021/395/2021/">providing harmonised terminology and methods, as well as practical guidelines and results allowing the intercomparison of continental or global Digital Elevation Models (DEM)</a>" (Strobl et al., 2021). Several publications have defined the framework of DEMIX, from <a href="https://doi.org/10.3390/rs13183581">the terminology and definitions</a> (Guth et al., 2021) to the <a href="https://doi.org/10.1109/TGRS.2024.3368015">DEM ranking methods</a> (Bielski et al., 2024). An additional methodology paper has been publicated regarding the assessment of <a href="https://doi.org/10.3390/ijgi13030096">planimetric displacements between DEMs</a> (Riazanoff et al., 2024), which are a common source of biases in DEM comparisons.</p> <h3>The DEMIX grid</h3> <p>Studies performed within the DEMIX framework rely on the <a href="../records/7504791">DEMIX grid</a> (Guth et al., 2023), a geodetic grid (EPSG:4326) dividing the world in areas of approximately 10x10km. These standard areas are called DEMIX tiles, and can be precisely located thanks to their identifier.</p> <h3>Criteria and scores</h3> <p>Within DEMIX, several criteria have been defined to assess the quality of DEMs. These criteria take as input a DEM and a DEMIX tile, and provide as output the score of the DEM for this specific tile. Repeating this process over several DEMs and DEMIX tiles of interest allow for a comparison of scores, leading to a ranking of DEMs. <a href="https://doi.org/10.1109/TGRS.2024.3368015">DEMIX rankings are based on the Randomized Complete Block Design (RCBD)</a> (Bielski et al., 2024).</p> <h2>This dataset</h2> <p>This dataset is composed of global maps of one map per (DEM, criterion) tuple. Each GeoTIFF map can be superimposed with the <a href="../records/7504791">DEMIX grid</a> (Guth et al., 2023) in a GIS (tested in QGIS 3.16).</p> <h3>Completeness criteria</h3> <p>The completeness criteria have originally been defined by Peter Strobl. A brief description of each criterion is given in the next table. Please see the column "Original document" and files of this repository for the complete definitions.</p> <table> <tbody> <tr> <td><strong>Criterion</strong></td> <td><strong>Description</strong></td> <td><strong>Requirements</strong></td> <td><strong> Original document</strong></td> </tr> <tr> <td>A01 - Product fractional cover</td> <td>Fraction of a DEMIX tile <strong>covered</strong> by the DEM product</td> <td>None</td> <td>See document "DEMIX_CDD-A01_20211103.docx"</td> </tr> <tr> <td>A02 - Valid data fraction</td> <td>Fraction of a DEMIX tile <strong>covered</strong> by <strong>valid </strong>pixels of the DEM product</td> <td>"No data" or "void" value in metadata</td> <td>See document "DEMIX_CDD-A02_20211103.docx"</td> </tr> <tr> <td>A03 - Primary data fraction</td> <td>Fraction of a DEMIX tile <strong>covered </strong>by <strong>valid </strong>pixels generated from the <strong>main source of data</strong> of the DEM product</td> <td>"No data" or "void" value in metadata + source data/editing mask</td> <td>See document "DEMIX_CDD-A03_20211103.docx"</td> </tr> <tr> <td>A04 - Valid land fraction</td> <td>Fraction of a DEMIX tile <strong>covered </strong>by <strong>valid </strong>pixels of <strong>land </strong>of the DEM product</td> <td>"No data" or "void" value in metadata + water body mask</td> <td>See document "DEMIX_CDD-A04_20211103.docx"</td> </tr> <tr> <td>A05 - Primary land fraction</td> <td>Fraction of a DEMIX tile <strong>covered </strong>by <strong>valid </strong>pixels of <strong>land </strong>generated from the <strong>main source of data </strong>of the DEM product</td> <td>"No data" or "void" value in metadata + water body mask + source data/editing mask</td> <td>See document "DEMIX_CDD-A05_20211103.docx"</td> </tr> </tbody> </table> <h3>DEMs and ancillary data</h3> <p>The following DEM products and ancillary layers have been used to generate the dataset.</p> <table> <tbody> <tr> <td><strong>Identifier</strong></td> <td><strong>Used layers</strong></td> <td><strong>Data access</strong></td> </tr> <tr> <td> <p>ASTGTM v003</p> </td> <td>ASTER GDEM elevations (dem.tif) + editing / source masks (num.tif)</td> <td><a href="https://lpdaac.usgs.gov/products/astgtmv003/">https://lpdaac.usgs.gov/products/astgtmv003/</a></td> </tr> <tr> <td> <p>ASTWBD v001</p> </td> <td>ASTER GDEM water body mask (att.tif)</td> <td><a href="https://lpdaac.usgs.gov/products/astwbdv001/">https://lpdaac.usgs.gov/products/astwbdv001/</a></td> </tr> <tr> <td> <p>AW3D30 v2003</p> </td> <td>ALOS World 3D elevations (DSM.tif) + editing / source / water body masks (MSK.tif)</td> <td><a href="https://www.eorc.jaxa.jp/ALOS/en/dataset/aw3d30/aw3d30_e.htm">https://www.eorc.jaxa.jp/ALOS/en/dataset/aw3d30/aw3d30_e.htm</a></td> </tr> <tr> <td>COP-DEM_GLO-30-DGED v2019_1</td> <td>Copernicus DEM GLO-30 elevations (DEM.tif) + editing (EDM.tif) + source (SRC.tif) + water body (WBM.tif) masks</td> <td><a href="https://spacedata.copernicus.eu/collections/copernicus-digital-elevation-model">https://spacedata.copernicus.eu/collections/copernicus-digital-elevation-model</a></td> </tr> <tr> <td>COP-DEM_GLO-90-DGED v2019_1</td> <td>Copernicus DEM GLO-90 elevations (DEM.tif) + editing (EDM.tif) + source (SRC.tif) + water body (WBM.tif) masks</td> <td><a href="https://spacedata.copernicus.eu/collections/copernicus-digital-elevation-model">https://spacedata.copernicus.eu/collections/copernicus-digital-elevation-model</a></td> </tr> <tr> <td> <p>NASADEM_HGT v001</p> </td> <td>NASADEM elevations (.hgt) + editing / source (.num) + water body (.swb) masks</td> <td><a href="https://lpdaac.usgs.gov/products/nasadem_hgtv001/">https://lpdaac.usgs.gov/products/nasadem_hgtv001/</a></td> </tr> <tr> <td> <p>SRTMGL1 v003</p> </td> <td>SRTMGL1 elevations (.hgt)</td> <td><a href="https://lpdaac.usgs.gov/products/srtmgl1v003/">https://lpdaac.usgs.gov/products/srtmgl1v003/</a></td> </tr> <tr> <td> <p>SRTMGL1N v003</p> </td> <td>SRTMGL1 editing / source / water body masks (.num)</td> <td><a href="https://lpdaac.usgs.gov/products/srtmgl1nv003/">https://lpdaac.usgs.gov/products/srtmgl1nv003/</a></td> </tr> </tbody> </table> <h3>Computation of scores</h3> <p>For each DEMIX tile and DEM, each "fractional cover" has been computed using the following procedure:</p> <ol> <li><strong>Crop DEMIX tile layers</strong> - The tiles of each DEM layer (elevations, editing, sources and water bodies) are cropped to the extent of the DEMIX tile.</li> <li><strong>Compute standardized layers</strong><strong> </strong>- Given the cropped DEM layers, four standardized layers are produced, which are: <ul> <li>Heights layer - Containing the heights of the DEM</li> <li>Land/water mask layer - Indicating whether DEM pixels are land or water: <ul> <li>0 = NO_DATA</li> <li>1 = BACKGROUND</li> <li>2 = INVALID</li> <li>3 = WATER</li> <li>4 = LAND</li> </ul> </li> <li>Source mask layer - Indicating the source data of DEM heights (or "edited" value): <ul> <li>0 = NO_DATA</li> <li>1 = BACKGROUND</li> <li>2 = INVALID</li> <li>3 = PRIMARY_DATA</li> <li>4 = EXTERNAL_DATA</li> <li>5 = EDITED</li> </ul> </li> <li>Valid mask layer - Indicating if the DEM pixels are valid or not: <ul> <li>0 = NO_DATA</li> <li>1 = BACKGROUND</li> <li>2 = INVALID</li> <li>3 = VALID</li> </ul> </li> </ul> </li> <li><strong>Retrieve pixel number N</strong><em><strong> </strong>-<strong> </strong></em>The total pixel number N is computed for one of the layers (all layers have the same number of pixels).</li> <li><strong>Retrieve criterion pixel number C </strong>-<strong> </strong>The criterion pixel number C is computed based on the standard layers, more precisely: <ul> <li>A01 - Product fractional cover - Number of pixels of <strong>valid mask layer equal to 1, 2 or 3</strong></li> <li>A02 - Valid data fraction - Number of pixels of <strong>valid mask layer equal to 3</strong></li> <li>A03 - Primary data fraction - Number of pixels of <strong>source mask layer equal to 3</strong></li> <li>A04 - Valid land fraction - Number of pixels of <strong>land/water mask layer equal to 4</strong></li> <li>A05 - Primary land fraction - Number of pixels of <strong>source mask layer equal to 3</strong> and<strong> land/water mask layer equal to 4</strong></li> </ul> </li> <li><strong>Compute the final score S</strong><strong> </strong>- The final score S is expressed as the following percentage: <strong>S = ceil(C/N*100)</strong></li> </ol> <h2>Known issues</h2> <p>The "SRTMGL1N v003" is known to have "tile repeating issues", where part of the data is wrongly flagged as water. This issue has been reported with no particular response from the providers of the DEM (see <a href="https://forum.earthdata.nasa.gov/viewtopic.php?t=2752">https://forum.earthdata.nasa.gov/viewtopic.php?t=2752</a>).</p> <p><strong>References:</strong></p> <ul> <li>Guth, P.L.; Strobl, P.; Gross, K.; Riazanoff, S. <em>DEMIX 10k Tile Data Set (1.0)</em> [Data set]. Zenodo 2023. <a href="https://doi.org/10.5281/zenodo.7504791" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.7504791</a></li> <li>Guth, P.L.; Van Niekerk, A.; Grohmann, C.H.; Muller, J.-P.; Hawker, L.; Florinsky, I.V.; Gesch, D.; Reuter, H.I.; Herrera-Cruz, V.; Riazanoff, S.; López-Vázquez, C.; Carabajal, C.C.; Albinet, C.; Strobl, P. <em>Digital Elevation Models: Terminology and Definitions</em>. Remote Sens. 2021, 13, 3581. <a href="https://doi.org/10.3390/rs13183581">https://doi.org/10.3390/rs13183581</a></li> <li>Riazanoff, S.; Corseaux, A.; Albinet, C.; Strobl, P.A.; López-Vázquez, C.; Guth, P.L.; Tadono, T. <em>Best BiCubic Method to Compute the Planimetric Misregistration between Images with Sub-Pixel Accuracy: Application to Digital Elevation Models</em>. <em>ISPRS Int. J. Geo-Inf.</em> 2024, <em>13</em>, 96. <a href="https://doi.org/10.3390/ijgi13030096">https://doi.org/10.3390/ijgi13030096</a></li> <li>Bielski, C.; López-Vázquez, C.; Grohmann, C.H.; Guth, P.L.; Hawker, L.; Gesch, D.; Trevisani, S.; Herrera-Cruz, V.; Riazanoff, S.; Corseaux, A.; Reuter, H.I.; Strobl, P.A.; <em>Novel Approach for Ranking DEMs: Copernicus DEM Improves One Arc Second Open Global Topography</em> in <em>IEEE Transactions on Geoscience and Remote Sensing</em>, vol. 62, pp. 1-22, 2024, Art no. 4503922. <a href="https://doi.org/10.1109/TGRS.2024.3368015">https://doi.org/10.1109/TGRS.2024.3368015</a></li> <li>Strobl, P.A.; Bielski, C.; Guth, P.L.; Grohmann, C.H.; Muller, J.P.; López-Vázquez, C.; Gesch, D.B.; Amatulli, G.; Riazanoff, S.; Carabajal, C. The Digital Elevation Model Intercomparison eXperiment DEMIX, a community based approach at global DEM benchmarking. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2021, XLIII-B4-2021, 395–400. <a href="https://doi.org/10.5194/isprs-archives-XLIII-B4-2021-395-2021">https://doi.org/10.5194/isprs-archives-XLIII-B4-2021-395-2021</a></li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.