Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

624

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

624 results for “Gold”

Learn how ShareScore rates datasets ↗
zenodo40/100

Fig. 16 in Solomon's Gold Mine: Description or redescription of 24 species of Caridina (Crustacea: Decapoda: Atyidae) freshwater shrimps from the Solomon Islands, including 11 new species

Fig. 16. Caridina tupaia de Mazancourt, Marquet & Keith, 2019. Cephalothorax. MNHN-IU-2018 -2884.

opencc-by-4.0Aug 2020View details →
zenodo40/100

Fig. 17 in Solomon's Gold Mine: Description or redescription of 24 species of Caridina (Crustacea: Decapoda: Atyidae) freshwater shrimps from the Solomon Islands, including 11 new species

Fig. 17. Caridina maeana sp. nov. a. First pereiopod. b. Second pereiopod. c. Third pereiopod. d. Fifth pereiopod. e. Dactylus of third pereiopod. f. Dactylus of fifth pereiopod. g. Pre-anal carina. h. Uropodal diaeresis. i. Telson. j. First male pleopod. k. Second male pleopod. l. Eggs. m. Cephalothorax. MNHN-IU-2018-2889: (a–i, m), MNHN-IU-2018-2891 (j–k) and MNHN-IU-2018-2895 (l).

opencc-by-4.0Aug 2020View details →
zenodo40/100

Fig. 12. Caridina buehleri Roux, 1934. a. First pereiopod. b. Second pereiopod. c. Third pereiopod. d. Fifth pereiopod. e in Solomon's Gold Mine: Description or redescription of 24 species of Caridina (Crustacea: Decapoda: Atyidae) freshwater shrimps from the Solomon Islands, including 11 new species

Fig. 12. Caridina buehleri Roux, 1934. a. First pereiopod. b. Second pereiopod. c. Third pereiopod. d. Fifth pereiopod. e. Dactylus of third pereiopod. f. Dactylus of fifth pereiopod. g. Pre-anal carina. h. Uropodal diaeresis. i. Telson. j. Pre-anal carina. k. Eggs. l. Cephalothorax. MNHN-IU-2018-2845 (a–i, k), MNHN-IU-2015-20 (j) and MNHN-IU-2018-2846 (l).

opencc-by-4.0Aug 2020View details →
zenodo40/100

Fig. 11 in Solomon's Gold Mine: Description or redescription of 24 species of Caridina (Crustacea: Decapoda: Atyidae) freshwater shrimps from the Solomon Islands, including 11 new species

Fig. 11. Caridina turipi sp. nov. a. First pereiopod. b. Second pereiopod. c. Third pereiopod. d. Fifth pereiopod. e. Dactylus of third pereiopod. f. Dactylus of fifth pereiopod. g. Pre-anal carina. h. Uropodal diaeresis. i. Telson. j. Eggs. k. Cephalothorax. MNHN-IU-2014-20883 (a–f, h, j–k), MNHN-IU-2014-20882 (g) and MNHN-IU-2014-20878 (i).

opencc-by-4.0Aug 2020View details →
zenodo40/100

Fig. 10. Caridina typus H. Milne Edwards, 1837. a. First pereiopod. b. Second pereiopod. c. Third pereiopod. d. Fifth pereiopod. e in Solomon's Gold Mine: Description or redescription of 24 species of Caridina (Crustacea: Decapoda: Atyidae) freshwater shrimps from the Solomon Islands, including 11 new species

Fig. 10. Caridina typus H. Milne Edwards, 1837. a. First pereiopod. b. Second pereiopod. c. Third pereiopod. d. Fifth pereiopod. e. Dactylus of third pereiopod. f. Dactylus of fifth pereiopod. g. Preanal carina. h. Uropodal diaeresis. i. Telson. j. First male pleopod. k. Second male pleopod. l. Eggs. m. Cephalothorax. MNHN-IU-2018-2825 (a–c, e, h), MNHN-IU-2018-2842 (d, f), MNHN-IU-2018 -2826 (g, j–k), MNHN-IU-2018-2833 (i), MNHN-IU-2018-2832 (l) and MNHN-IU-2018-2824 (m).

opencc-by-4.0Aug 2020View details →
zenodo40/100

Fig. 8. Caridina mertoni Roux, 1911. a. First pereiopod. b. Second pereiopod. c. Third pereiopod. d. Fifth pereiopod. e in Solomon's Gold Mine: Description or redescription of 24 species of Caridina (Crustacea: Decapoda: Atyidae) freshwater shrimps from the Solomon Islands, including 11 new species

Fig. 8. Caridina mertoni Roux, 1911. a. First pereiopod. b. Second pereiopod. c. Third pereiopod. d. Fifth pereiopod. e. Dactylus of third pereiopod. f. Dactylus of fifth pereiopod. g. Pre-anal carina. h. Uropodal diaeresis. i. Telson. j. First male pleopod. k. Second male pleopod. l. Eggs. m. Cephalothorax. n–o. Rostrum variations. MNHN-IU-2018-2817 (a–f, m), MNHN-IU-2018-2815 (h–i, l), MNHN-IU-2018-2816 (j–k), MNHN-IU-2017-2109 (n) and MNHN-IU-2018-2819 (g, o).

opencc-by-4.0Aug 2020View details →
zenodo40/100

Fig. 3 in Solomon's Gold Mine: Description or redescription of 24 species of Caridina (Crustacea: Decapoda: Atyidae) freshwater shrimps from the Solomon Islands, including 11 new species

Fig. 3. Caridina barakoma sp. nov. a. First pereiopod. b. Second pereiopod. c. Third pereiopod. d. Fifth pereiopod. e. Dactylus of third pereiopod. f. Dactylus of fifth pereiopod. g. Pre-anal carina. h. Uropodal diaeresis. i. Telson. j. First male pleopod. k. Second male pleopod. l. Eggs. m. Cephalothorax. MNHN-IU-2014-20807 (a–i, l–m) and MNHN-IU-2014-20805 (j–k).

opencc-by-4.0Aug 2020View details →
zenodo40/100

Fig. 6 in Solomon's Gold Mine: Description or redescription of 24 species of Caridina (Crustacea: Decapoda: Atyidae) freshwater shrimps from the Solomon Islands, including 11 new species

Fig. 6. Caridina choiseul sp. nov. a. First pereiopod. b. Second pereiopod. c. Third pereiopod. d. Fifth pereiopod. e. Dactylus of third pereiopod. f. Dactylus of fifth pereiopod. g. Pre-anal carina. h. Uropodal diaeresis. i. Telson. j. First male pleopod. k. Second male pleopod. l. Eggs. m–n. Rostrum variations. o. Cephalothorax. MNHN-IU-2014-20833 (a–g (top, with spine), h, l), MNHN-IU-2014-20836 (g (botom, without spine)), MNHN-IU-2014-20831 (i), MNHN-IU-2014-20829 (j–k), MNHN-IU-2014-20827 (m), MNHN-IU-2014-20837 (n) and MNHN-IU-2014-20826 (o).

opencc-by-4.0Aug 2020View details →
zenodo40/100

PharmaCoNER corpus: gold standard annotations of Pharmacological Substances, Compounds and proteins in Spanish clinical case reports

<p><strong>Intro:</strong></p><p>The PharmaCoNER corpus (divided into train, dev and test) is a Gold Standard manually annotated dataset used for the the PharmaCoNER shared task posed at BIONLP-ST (at EMNLP). In addition, we include here the PharmaCoNER background set. It contains the train, development and test sets of the two subtasks (subtask-1 and subtask-2) with Gold Standard annotations. In addition, it contains the documents of the background set, without annotations.</p><p>The PharmaCoNER corpus consists of:</p><ul><li>Manually classified clinical case sections derived from Open access Spanish medical publications, named the Spanish Clinical Case Corpus (SPACCC).</li><li>It was manually selected by a practicing oncologist and revised by a clinical documentalist to assure that records were relevant/representative and resembled structure and content relevant to process clinical records.</li><li>The final corpus: 1000 clinical cases 16,504 sentences</li><li>The corpus contains a total of 396,988 words, with an average of 396.2 words per clinical case.</li><li>It covers a range of medical disciplines including oncology, urology, cardiology, pneumology or infections diseases, etc.</li><li>The corpus has been annotated at the mention level by experts in medicinal chemistry and pharmacology following a granular annotation scheme covering four mention types:<ul><li><i>Entity type 1 (NORMALIZABLES)</i>: mentions of chemicals that can be manually normalized to a unique concept identifier (primarily SNOMED-CT).</li><li><i>Entity type 2 (NO_NORMALIZABLES)</i>: mentions of chemicals that could not be normalized manually to a unique concept identifier.</li><li><i>Entity type 3 (PROTEINAS)</i>: mentions of proteins/genes following an adaptation of the BioCreative GPRO track annotation guidelines (includes peptides, peptide hormones &amp; antibodies).</li><li><i>Entity type 4 (UNCLEAR )</i>: cases of general substance class mentions of clinical relevance, including certain pharmaceutical formulations, general treatments, chemotherapy programs, and vaccines.</li><li>Mentions class "<i>UNCLEAR</i>" (not evaluated for the PharmaCoNER track)</li></ul></li></ul><p>&nbsp;</p><p>&nbsp;</p><p><strong>Please, cite:&nbsp;</strong></p><p>A. G. Agirre, M. Marimon, A. Intxaurrondo, O. Rabal, M. Villegas, M. Krallinger, Pharmaconer: Pharmacological substances, compounds and proteins named entity recognition track, in: Proceedings of The 5th Workshop on BioNLP Open Shared Tasks, 2019, pp. 1–10.</p><p>&nbsp;</p><p><strong>Annotation quality</strong></p><p>Inter-annotator agreement: 93% for annotation, 73% for mapping.</p><p>For more information, see the <a href="https://paperswithcode.com/paper/pharmaconer-pharmacological-substances">paper</a>.</p><p>&nbsp;</p><p><strong>Format</strong></p><p>For subtask 1 annotations are distributed in <i>Brat</i> format. (More info at Brat webpage&nbsp;https://brat.nlplab.org/standoff.html)</p><p>For subtask-2, codes are associated with each document are given in a <i>TSV</i> file with the following columns:&nbsp;</p><blockquote><p>filename&nbsp;&nbsp; &nbsp;code</p></blockquote><p>&nbsp;</p><p><strong>Shared task goal:</strong></p><p>In the two subtasks, the goal is to predict the annotations of the test files (either the ANN files or the TSV with the codes) given only the plain text files.&nbsp;</p><p>&nbsp;</p><p><strong>Resources:</strong></p><ul><li><a href="https://temu.bsc.es/pharmaconer/"><strong>Web</strong></a></li><li><a href="https://www.aclweb.org/anthology/D19-5701.pdf"><strong>Citation</strong></a><strong>:&nbsp;</strong>A. G. Agirre, M. Marimon, A. Intxaurrondo, O. Rabal, M. Villegas, M. Krallinger, Pharmaconer: Pharmacological substances, compounds and proteins named entity recognition track, in: Proceedings of The 5th Workshop on BioNLP Open Shared Tasks, 2019, pp. 1–10.</li><li><a href="https://doi.org/10.5281/zenodo.4271908"><strong>Silver Standard corpus</strong></a></li><li><a href="https://doi.org/10.5281/zenodo.3763276"><strong>Annotation guidelines</strong></a></li><li><a href="https://github.com/TeMU-BSC/PharmaCoNER-Tagger"><strong>PharmaCoNER tagger</strong></a></li><li><a href="https://www.youtube.com/watch?v=B3ZzJl5OMkY"><strong>Youtube video(general setting)</strong></a></li><li><a href="https://www.slideshare.net/MartinKrallinger/pharmaconer-pharmacological-substances-compounds-and-proteins-named-entity-recognition-track-at-bionlpost-workshop-november-4-skycity-rm-2-hong-kong-emnlp2019"><strong>Slides PharmacoNER overview talk at BIONLP-ST / EMNLP&nbsp;</strong></a></li></ul><p>For further information, please visit <a href="https://temu.bsc.es/pharmaconer/">https://temu.bsc.es/pharmaconer/</a> or email us at encargo-pln-life@bsc.es</p><p>Copyright (c) 2018 Secretaría de Estado para el Avance Digital (SEAD)</p><p>&nbsp;</p><p><strong>License</strong></p><p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p><p>&nbsp;</p><p><strong>Contact</strong></p><p>If you have any questions or suggestions, please contact us at:</p><p><br>- Martin Krallinger (&lt;krallinger [dot] martin [at] gmail [dot] com&gt;)</p><p><strong>Additional resources and corpora</strong></p><p>If you are interested in PharmaCoNER, you might want to check out these corpora and resources:</p><ul><li><a href="https://zenodo.org/records/7614764">DisTEMIST</a> (Corpus of disease mentions and normalization to SNOMED CT, same document collection)</li><li><a href="https://zenodo.org/records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT, same document collection)</li><li><a href="https://zenodo.org/records/8413866">SympTEMIST</a> (Corpus of symptoms, signs and findings mentions and normalization, same document collection)</li><li><a href="https://zenodo.org/records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI), modified synthetic verions of the document collection)</li><li><a href="https://zenodo.org/records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization, different document collection)</li><li><a href="https://zenodo.org/records/3837305">CodiESp</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version, same document collection)</li><li><a href="https://zenodo.org/records/7684093">LivingNER</a> (Corpus of mentions of species, including human/family members, pathogens, food, etc.. and normalization to NCBI Taxonomy, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags, same document collection)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries), same document collection)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags, same document collection)</li><li><a href="https://zenodo.org/records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, clinical trials, and patent abstracts, different document collection)</li></ul>

opencc-by-4.0Nov 2020View details →
zenodo40/100

MEDDOCAN corpus: gold standard annotations for Medical Document Anonymization on Spanish clinical case reports

<p><strong>Intro:</strong></p> <p>Meddocan shared task dataset (divided in train, dev and test). In addition, we include here the Meddocan background set.</p> <p>It contains the training, development and test sets of the Meddocan shared task with Gold Standard annotations.</p> <p>In addition, it contains the documents of the background set, without annotations.</p> <p>&nbsp;</p> <p><strong>Annotation quality</strong></p> <p>Inter-annotator agreement: 98%&nbsp;</p> <p>For more information, see the <a href="http://ceur-ws.org/Vol-2421/MEDDOCAN_overview.pdf">paper</a>.&nbsp;</p> <p>&nbsp;</p> <p><strong>Format:</strong></p> <p>Annotations are distributed in Brat format. See&nbsp;<a href="https://brat.nlplab.org/standoff.html">Brat webpage</a>&nbsp;for more information.</p> <p>In addition, annotations are also distributed in XML format (based on i2b2 XML format).</p> <p>In the <a href="https://temu.bsc.es/meddocan/index.php/resources/">Meddocan webpage</a>, there is a script to convert&nbsp;between MEDDOCAN-Brat, MEDDOCAN-XML, and i2b2 formats.</p> <p>&nbsp;</p> <p><strong>Shared task goal:</strong></p> <p>In the three subtasks, the goal will be to predict the annotations&nbsp;given only the plain text files.&nbsp;</p> <p>&nbsp;</p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/meddocan/">Web</a></strong></li> <li><strong>Citation:&nbsp;</strong>Montserrat Marimon et al. &ldquo;Automatic De-identification of Medical Texts in Spanish: the MEDDOCAN Track, Corpus, Guidelines, Methods and Evaluation of Results.&rdquo; In: IberLEF@ SEPLN. 2019, pp. 618&ndash;638.</li> <li><strong>Silver Standard corpus</strong></li> <li><a href="https://doi.org/10.5281/zenodo.4279337"><strong>Annotation guidelines</strong></a></li> </ul> <p>&nbsp;</p> <p>For further information, please visit&nbsp;<a href="https://temu.bsc.es/meddocan/">https://temu.bsc.es/meddocan/</a>&nbsp;or email us at encargo-pln-life@bsc.es</p> <p>Copyright (c) 2019 Secretar&iacute;a de Estado para el Avance Digital (SEAD)</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Topology and structure of Au144(SRNH3+)60 from "Atomistic Simulations of Functional Au144(SR)60 Gold Nanoparticles in Aqueous Environment"

<p>Positively&nbsp;charged monolayer-protected gold nanoparticles (AuNPs) structure&nbsp;and topology files for GROMACS&nbsp;used in DOI:&nbsp;10.1021/jp301094m. &nbsp;The final structure of&nbsp;the simulation&nbsp;reported in&nbsp;DOI: 10.1021/jp301094m&nbsp;for the neutral case is&nbsp;provided.</p> <p>The&nbsp;gold nanoparticle contain a core of 144&nbsp;Au&nbsp;atoms and&nbsp;60&nbsp;functionalized alkanethiol side groups (undecanyl chain, R = C11H22), each possessing a positively charged amonium&nbsp;terminal group.</p> <p>When using this structure do not forget to cite&nbsp;DOI:&nbsp;10.1021/jp301094m.&nbsp;</p> <p>NOTE1: Different versions for the topology files are provided of both&nbsp;AuNPs. All versions were used for the publication. The changes only affect the core surface and therefore had no influence in the reported properties. Still we recommentd using the latest version:&nbsp;AU144SRNH360_v3.itp.</p> <p>NOTE2: The original simulations used GROMACS&nbsp;4.0.5. The files should work, however, as well up to GROMACS&nbsp;4.6.7.</p>

opencc-by-4.0Apr 2012View details →
zenodo40/100

Topology and structure of Au144(SRCOO-)60 from "Atomistic Simulations of Functional Au144(SR)60 Gold Nanoparticles in Aqueous Environment"

<p>Negatively charged monolayer-protected gold nanoparticles (AuNPs) structure&nbsp;and topology files for GROMACS&nbsp;used in DOI:&nbsp;10.1021/jp301094m. &nbsp;The final structure of&nbsp;the simulation&nbsp;reported in&nbsp;DOI: 10.1021/jp301094m&nbsp;for the neutral case is&nbsp;provided.</p> <p>The&nbsp;gold nanoparticle contain a core of 144&nbsp;Au&nbsp;atoms and 60&nbsp;functionalized alkanethiol side groups (undecanyl chain, R = C11H22), each possessing a negatively charged carboxylic&nbsp;terminal group.</p> <p>When using this structure do not forget to cite DOI:&nbsp;10.1021/jp301094m.&nbsp;</p> <p>NOTE1: Different versions for the topology files are provided of both&nbsp;AuNPs. All versions were used for the publication. The changes only affect the core surface and therefore had no influence in the reported properties. Still we recommentd using the latest version:&nbsp;AU144SRCOO60_v2.itp.</p> <p>NOTE2: The original simulations used GROMACS&nbsp;4.0.5. The files should work, however, as well up to GROMACS&nbsp;4.6.7.</p>

opencc-by-4.0Apr 2012View details →
zenodo40/100

Supplementary Material: Computational Study of Quasi-2D Liquid State in Free Standing Platinum, Silver, Gold, and Copper Monolayers

<p>Supplementary files for&nbsp;<em>Condensed Matter</em>&nbsp;<strong>2016</strong>, <em>1</em>(1), 1; doi:10.3390/condmat1010001;&nbsp;http://www.mdpi.com/&nbsp;2410-3896/1/1/1.</p> <p>Captions:</p> <p><strong>Video S1.</strong>&nbsp;(Pt 2400 K 5 ps)&nbsp;5 ps Molecular Dynamics Movie of Pt Freestanding Monolayer at 2400 K.&nbsp;</p> <p><strong>Video S2.</strong>&nbsp; (Ag 1050K 6 ps)&nbsp;6 ps Molecular Dynamics Movie of Ag Freestanding Monolayer at 1050 K.<br /> <br /> <strong>Video S3.</strong>&nbsp;(Au 1600K 4ps)&nbsp;4 ps Molecular Dynamics Movie of Au Freestanding Monolayer at 1600 K.<br /> <br /> <strong>Video S4.</strong>&nbsp;(Cu 1400K 3ps)&nbsp;3 ps Molecular Dynamics Movie of Cu Freestanding Monolayer at 1400 K.&nbsp;</p>

opencc-by-4.0Mar 2016View details →
zenodo40/100

Medical Relation Extraction Gold Standard with CrowdTruth

<p>The lack of annotated datasets for training and benchmarking is one of the main challenges of Clinical Natural Language Processing. In addition, current methods for collecting annotations attempt to minimize disagreement between annotators, and therefore fail to model the ambiguity inherent in language. We propose the <strong>CrowdTruth</strong> method for collecting medical ground truth through crowdsourcing, based on the observation that disagreement between annotators can signal ambiguity in the text, target semantics, or the worker&#39;s interpretation.</p> <p>This repository contains a dataset of 3,984 English sentences for medical relation extraction, centering on the cause and treat medical relations, that have been processed with CrowdTruth disagreement analytics to capture ambiguity. In addition, we provide the raw crowdsourcing data used to compile this ground truth, as well as the task templates used to collect the data on CrowdFlower.</p>

opencc-by-sa-4.0Apr 2016View details →
zenodo40/100

A rule based Tibetan part-of-speech (POS) tagger for the creation of gold standard training data

<p>This rule based Tibetan part-of-speech (POS) tagger was prepared in the course of the research project &#39;Tibetan in Digital Communication&#39; (2012-2015) hosted at SOAS, University of London and funded by the UK&#39;s Arts and Humanities Research Council (grant code: AH/J00152X/1). For a description of the tag set see Garrett et al. 2014. and Garrett et al. 2015. For a description of the tagger itself see Garrett et al. 2014. Note that the tagger must be used together with a lexicon (for example Hill &amp; Garrett 2017a). One must use one&#39;s own script to tag all words with all tags in the lexicon and then apply the tagger to remove incorrect tags.</p> <p>On the associated corpus of 318,230 words (Hill &amp; Garrett 2017b) the lexical tagger (i.e. simply applying all available tags to all words) tags 141,911 words with the correct unique tag, achieves as accuracy of 1.000 (by definition getting the right tag among others for each word) with an ambiguity of 2.73111. In contrast, the Rule Tagger tags 241,256 words with the correct unique tag, achieves an accuracy of 0.99893 and an ambiguity of 1.38577.</p> <p>Because this tagger does not achieve ambiguity 1.000 it is not suitable for tagging large scale corpora, but instead is useful for the creation of gold standard training data.</p> <p>N.B. In some rare cases the tagger removes all POS-tags.</p>

opencc-by-4.0May 2017View details →
zenodo40/100

Documentation of a Pyu inscription on a gold ring (PYU105) held by the Śrī Kṣetra museum

<p>This data set includes photographs (.jpg) and related files documenting a Pyu inscription (inventory number PYU105) on a gold ring held at the Śrī Kṣetra museum. The photographer was James Miles or Archeovision, working on behalf of the Pyu epigraphy sub-project (PI, Nathan W. Hill of SOAS University of London) of the ERC synergy grant "Beyond Boundaries: Religion, Region, Language and the State" (Identifier: ASIA 609823) in collaboration with the project "From Vijayapuri to Sriksetra? The Beginnings of Buddhist Exchange across the Bay of Bengal as Witnessed by Inscriptions from Andhra Pradesh and Myanmar" (PI Arlo Giffiths of the EFEO) funded by The Robert H. N. Ho Family Foundation.</p>

opencc-by-4.0Nov 2016View details →
zenodo40/100

Files and code for English dictionaries, gold and silver standard corpora for biomedical natural language processing related to SARS-CoV-2 and COVID-19

<p><span lang="EN-GB">Automated information extraction with natural language processing (NLP) tools is required to gain systematic insights from the large number of COVID-19 publications, reports and social media posts, which far exceed human processing capabilities.&nbsp;</span></p> <p><span lang="EN-GB">Here we present an NLP toolbox comprising COVID-19-related dictionaries and annotated corpora in English as well as useful code and workflows for their update and use. The dictionaries contain terms referring to the COVID-19 disease, the SARS-CoV-2 virus, its variants and common mutations, respectively. They were used together with the EasyNER NLP tool to extract and annotate all 764&nbsp;398 abstracts in the CORD-19 dataset, creating a very large silver standard corpus (named Lund-Annotated-CORD-19 corpus). This was complemented with a small gold standard corpus consisting of PubMed abstracts manually annotated for key entity classes such as disease, virus, symptom, protein/gene, cell type, chemical and species terms. </span></p> <p><span lang="EN-GB">The toolbox can support various text analysis tasks related to COVID-19 such as named entity recognition and co-mention analysis. A preliminary version of the toolbox, which was released early in the pandemic, was</span><span lang="EN-GB"> for example already used to create a COVID-19 knowledge graph and study the evolution and variation of COVID-19-related terminology. In addition, the toolbox can be applied in the development of other NLP tools, for example to train and evaluate large language models.</span></p> <p><span lang="EN-GB">When using the toolbox, please cite this record and the associated article.</span></p> <p>&nbsp;</p> <p>&nbsp;</p>

openJun 2022View details →
dryad40/100

Twisted epitaxy of gold nanodiscs grown between twisted substrate layers of molybdenum disulfide

<p>We expand the concept of epitaxy to a regime of "twisted epitaxy" with the epilayer crystal orientation between two substrates influenced by their relative orientation. We annealed nanometer-thick gold (Au) nanoparticles between two substrates of exfoliated hexagonal molybdenum disulfide (MoS<sub>2</sub>) with varying orientation of their basal planes with a mutual twist angle from 0° to 60°. Transmission electron microscopy studies show that the Au alignment is midway between that of the top and bottom MoS<sub>2</sub> when the twist angle of the bilayer is small (&lt; ~7°). For larger twist angles, Au has only a small misorientation with the bottom MoS<sub>2</sub> that varies approximately sinusoidally with the twist angle of the bilayer MoS<sub>2</sub>. Four-dimensional scanning transmission electron microscopy analysis further reveals a periodic strain variation (&lt; |±0.5%|) in the Au nanodiscs associated with the twisted epitaxy, consistent with the Moiré registry of the two MoS<sub>2</sub> twisted layers.</p>

opencc-zeroNov 2023View details →
zenodo40/100

GeoEDdA: A Gold Standard Dataset for Named Entity Recognition and Span Categorization Annotations of Diderot & d'Alembert's Encyclopédie

<p>This repository contains a gold standard dataset for named entity recognition and span categorization annotations from Diderot &amp; d&rsquo;Alembert&rsquo;s Encyclop&eacute;die entries.</p> <p>The dataset is available in the following formats:</p> <ul> <li>JSONL format provided by <a href="https://prodi.gy/" rel="nofollow">Prodigy</a></li> <li>binary spaCy format (ready to use with the spaCy train pipeline)</li> </ul> <p>The Gold Standard dataset is composed of 2,200 paragraphs out of 2,001 Encyclop&eacute;die's entries randomly selected. All paragraphs were written in 19th-century French.</p> <p>The spans/entities were labeled by the project team along with using pre-labelling with early machine learning models to speed up the labelling process. A train/val/test split was used. Validation and test sets are composed of 200 paragraphs each: 100 classified under 'G&eacute;ographie' and 100 from another knowledge domain. The datasets have the following breakdown of tokens and spans/entities.</p> <h2>Tagset</h2> <ul> <li><strong>NC-Spatial</strong>: a common noun that identifies a spatial entity (nominal spatial entity) including natural features, e.g. <code>ville</code>,&nbsp;<code>la rivi&egrave;re</code>, <code>royaume</code>.</li> <li><strong>NP-Spatial</strong>: a proper noun identifying the name of a place (spatial named entities), e.g. <code>France</code>, <code>Paris</code>, <code>la Chine</code>.</li> <li><strong>ENE-Spatial</strong>: nested spatial entity , e.g. <code>ville de France</code> , <code>royaume de Naples</code>, <code>la mer Baltique</code>.</li> <li><strong>Relation</strong>: spatial relation, e.g. <code>dans</code>, <code>sur</code>, <code>&agrave; 10 lieues de</code>.</li> <li><strong>Latlong</strong>: geographic coordinates, e.g. <code>Long. 19. 49. lat. 43. 55. 44.</code></li> <li><strong>NC-Person</strong>: a common noun that identifies a person (nominal spatial entity), e.g. <code>roi</code>, <code>l'empereur</code>, <code>les auteurs</code>.</li> <li><strong>NP-Person</strong>: a proper noun identifying the name of a person (person named entities), e.g. <code>Louis XIV</code>, <code>Pline</code>.</li> <li><strong>ENE-Person</strong>: nested people entity, e.g. <code>le czar Pierre</code>, <code>roi de Mac&eacute;doine</code>.</li> <li><strong>NP-Misc</strong>: a proper noun identifying entities not classified as spatial or person, e.g. <code>l'Eglise</code>, <code>1702</code>, <code>P&eacute;lasgique</code></li> <li><strong>ENE-Misc</strong>: nested named entity not classified as spatial or person, e.g. <code>l'ordre de S. Jacques</code>, <code>la d&eacute;claration du 21 Mars 1671</code>.</li> <li><strong>Head</strong>: entry name</li> <li><strong>Domain-Mark</strong>: words indicating the knowledge domain (usually after the head and between parenthesis), e.g. <code>G&eacute;ographie</code>, <code>Geog.</code>, <code>en Anatomie</code>.</li> </ul> <h2>HuggingFace</h2> <p>The GeoEDdA dataset is available on the HuggingFace Hub: <a href="https://huggingface.co/datasets/GEODE/GeoEDdA">https://huggingface.co/datasets/GEODE/GeoEDdA</a></p> <h2>spaCy Custom Spancat trained on Diderot &amp; d&rsquo;Alembert&rsquo;s Encyclop&eacute;die entries</h2> <p>This dataset was used to train and evaluate a custom spancat model for French using <a href="https://spacy.io/" rel="nofollow">spaCy</a>. The model is available on HuggingFace's model hub: <a href="https://huggingface.co/GEODE/fr_spacy_custom_spancat_edda" rel="nofollow">https://huggingface.co/GEODE/fr_spacy_custom_spancat_edda</a>.</p> <h2>Acknowledgement</h2> <p>The authors are grateful to the <a href="https://aslan.universite-lyon.fr/" rel="nofollow">ASLAN project</a> (ANR-10-LABX-0081) of the Universit&eacute; de Lyon, for its financial support within the French program "Investments for the Future" operated by the National Research Agency (ANR). Data courtesy the <a href="https://artfl-project.uchicago.edu/" rel="nofollow">ARTFL Encyclop&eacute;die Project</a>, University of Chicago.</p>

opencc-by-sa-4.0Jan 2024View details →
zenodo40/100

Determination of antibacterial and photothermal properties of novel composites based on graphene oxide/reduced graphene oxide, gold nanoparticles, and graphene quantum dots

<p>HR-TEM.zip - HR-TEM files, file type .jpg</p> <p>FTIR.zip - FTIR spectra, file type .spa</p> <p>Photoluminescence.zip - PL spectra, file type .opju</p> <p>UV-Vis.opju - Origin file with UV-Vis spectra combined</p> <p>Raman 532 nm.opju - Origin file with Raman spectra combined</p> <p>ABDA.opju - Origin file with singlet oxygen production measurements</p> <p>Contact angle.png - Image with contact angle values</p> <p>Antibacterial analysis.png - Image representing antibacterial growth inhibition analysis</p> <p>XRD.zip - XRD spectra, file type .dat</p>

opencc-by-4.0Oct 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record