Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7
datasets available to search
ShareScore release 0.9.0
Dataset results
7 results for “Biotea”
Biotea Linked Data for Open Access Full text PubMed Central
<p>this is the dataset from a-b</p>
biotea i-n dataset of rdf for pmc
<p>i-n dataset of rdf for pmc</p>
Biotea
<p>RDF for Pubmed Central</p>
biotea, RDf for pubmed central, c-h dataset
<p>C-H dataset of the RDF for pubmed central </p>
Biotea-2-Bioschemas test data
<p>Biotea-2-Bioschemas mapps Biotea model to schema.org following the approach proposed by Bioschemas. Here we present the test data used in Biotea GitHub pages, corresponding to 2596 PubMed Open Access (PMC-OA) subset publications together with the software used to render schema.org markup.</p> <p>Date deposited includes (i) publications retrieved from PMC-OA API, i.e., full text in JATS/XML, (ii) ontology terms recognized in the abstracts and obtained from the NCBO Annotator, i.e., semantic annotations, and (iii) the same annotations following the PubAnnotation format.</p> <p>Software deposited includes (i) biotea-bioschemas-metadata which parses JATS/XML files and creates Bioschemas markup including metadata, abstract and references, (ii) biotea-bioschemas-annotations which parses PubAnnotation annotations and creates Bioschemas markup, and (iii) biotea-bioschemas-showcase which uses the other two in order to display markup in a graphical basic way and render it as a script element in the HTML following the JSON-LD format. The corresonding GitHub repositories are: (i) https://github.com/biotea/biotea-bioschemas-metadata, (ii) https://github.com/biotea/biotea-bioschemas-annotations, and (iii) https://github.com/biotea/biotea-bioschemas-showcase.</p> <p>Biotea-2-bioschemas can be seen in action at http://biotea.github.io/bioschemas/</p>
Biotea dataset (vr. July 2012)
<p><strong>Background</strong></p> <p>Information reported by scientific literature still remains locked up in discrete documents that are not always interconnected or machine-readable. The Semantic Web together with approaches such as the Resource Description Framework (RDF) and the Linked Open Data (LOD) initiative offer a connectivity tissue that can be used to support the generation of self-describing, machine-readable documents.</p> <p> </p> <p><strong>Results</strong></p> <p>Biotea is an approach to generate RDF from scholarly documents. Our RDF model makes extensive use of existing ontologies and semantic enrichment services. Our dataset comprises 270,834 articles from PubMed Open Central in RDF/XML distributed in 404 zipped files. The RDFization process takes care of metadata, e.g., title, authors and journal, as well as semantic annotations on biological entities along the full text. Biological entities are extracted by using the NCBO Annotator and Whatizit.</p> <p>We use the Bibliographic Ontology (BIBO), Dublin Core Metadata Initiative Terms (DCMI-terms), and the Provenance Ontology (PROV-O) to model the bibliographic metadata. Links to related pages such as PubMed HTML articles are provided via rdfs:seeAlso while links to other semantic representation such as Bio2RDF PubMed articles are provided via owl:sameAs.</p> <p>The NCBO Annotator is used to extract entities covering ChEBI for chemicals; Pathway, and Functional Genomics Data Society (MGED) for genes and proteins; Master Drug Data Base (MDDB), NDDF, and NDFRT for drugs; SNOMED, SYMP, MedDRA, MeSH, MedlinePlus Health Topics (MedlinePlus), Online Mendelian Inheritance in Man (OMIM), FMA, ICD10, and Ontology for Biomedical Investigations (OBI) for diseases and medical terms; PO for plants; and MeSH, SNOMED, and NCIt for general terms.</p> <p>Whatizit is used for GO, UniProt proteins, UniProt Taxonomy, and diseases mapped to the UMLS; UniProt taxa are also mapped to NCBI Taxon vocabulary.</p> <p> </p> <p><strong>Conclusions</strong></p> <p>Biotea delivers models and tools for metadata enrichment and semantic processing of biomedical documents. Our dataset makes it easier to access to the first bunch of RDFized articles following the Biotea model. Our future plans include updating our dataset on regular basis in order to incorporate the latest articles added to the PubMed Open Central collection, next delivery is planned for the first half of 2017. Following datasets will support a mapping to the Semanticscience Integrated Ontology (SIO) in order to accomplish to the guidelines set by Bio2RDF.</p> <p> </p> <p><strong>Notes</strong></p> <p>Biotea approach in full is available at http://jbiomedsem.biomedcentral.com/articles/10.1186/2041-1480-4-S1-S5 (Garcia Castro, L.J., C. McLaughlin, and A. Garcia, <em>Biotea: RDFizing PubMed Central in Support for the Paper as an Interface to the Web of Data.</em> Biomedical semantics, 2013. <strong>4 Suppl 1</strong>: p. S5).</p> <p>Biotea algorithms are publicly available at https://github.com/biotea</p>
Biotea sample data
<p>This is the RDF loaded at http://biotea.linkeddata.es/sparql and described at http://biotea.github.io</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.