Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3
datasets available to search
ShareScore release 0.7.1
Dataset results
3 results for “ontology enrichment”
Processed proteomic and phosphoproteomic timeseries from Ostreococcus tauri, with Gene Ontology enrichment, from "A phospho-dawn of protein modification anticipates light onset in the picoeukaryote O. tauri"
<p>Diel regulation of protein levels and protein modification had been less studied than transcript rhythms. These data tables in .XLSX format report partial proteome (Table_S1) and phosphoproteome data (Table_S2), assayed using shotgun mass-spectrometry, from cultures of the alga <em>Ostreococcus tauri </em>under light-dark cycles, sampled at Zeitgeber times (ZT, hours) 0, 4, 8, 12, 16 and 20. 10% of quantified proteins but two-thirds of phosphoproteins were rhythmic. Gene Ontology enrichment analysis was applied to infer the functional enrichment of the proteins or phosphoproteins, grouped by their loadings in PCA analysis (Table_S3), by hierarchical clustering (Table_S4) or by the peak time of their rhythmic profile (Table_S5).Prompted by night-peaking and apparently dark-stable proteins, we also tested the proteome of cultures transferred to prolonged darkness for 24, 48, 72 or 96h (Table_S6), where the proteome changed less than under the diel cycle. The raw data are available from ProteomeXchange, with identifiers PXD001734, PXD001735 and PXD002909.</p>
Ontology Enrichment from Texts (OET): A Biomedical Dataset for Concept Discovery and Placement
<p>A biomedical dataset supporting ontology enrichment from texts, by concept discovery and placement, adapting the MedMentions dataset (PubMed abstracts) with SNOMED CT of versions in 2014 and 2017 under the Diseases (disorder) sub-category and the broader categories of Clinical finding, Procedure, and Pharmaceutical / biologic (CPP) product.</p> <p>The dataset is documented in the work, <em>Ontology Enrichment from Texts: A Biomedical Dataset for Concept Discovery and Placement</em>, on arXiv: <a href="https://arxiv.org/abs/2306.14704">https://arxiv.org/abs/2306.14704</a> (CIKM 2023). The companion code is available at https://github.com/KRR-Oxford/OET.</p> <p>Out-of-KB mention discovery (including the settings of mention-level data) is further partly documented in the work, <em>Reveal the Unknown: Out-of-Knowledge-Base Mention Discovery with Entity Linking</em>, on arXiv: <a href="https://arxiv.org/abs/2302.07189">https://arxiv.org/abs/2302.07189</a> (CIKM 2023).</p> <p>ver4: we made a version of mention-level data for out-of-KB discovery and concept placement separately: the former (for out-of-KB discovery) has out-of-KB mentions in training data, while the latter (for concept placement) has only out-of-KB mentions during the evaluation (validation and test) and not in the training data. Also, we split the original "test-NIL.jsonl" (now "test-NIL-all.jsonl") into "valid-NIL.jsonl" and "test-NIL.jsonl" for a better evaluation.</p> <p>ver3: we revised and updated mention-level data (syn_full, synonym augmentation setting) and the folder structure, and also updated the edge catalogues with complex edges.</p> <p>ver2: we revised the mention-level data by only keeping out-of-KB mentions (or "NIL" mentions) associated with one-hop edges (including leaf nodes, as <leaf node, NULL>) and two-hop edges in the ontology (SNOMED CT 20140901).</p> <p>Acknowledgement of data sources and tools below:</p> <p>* SNOMED CT https://www.nlm.nih.gov/healthit/snomedct/archive.html (and use snomed-owl-toolkit to form .owl files)<br>* UMLS https://www.nlm.nih.gov/research/umls/licensedcontent/umlsarchives04.html (and mainly use MRCONSO for mapping UMLS to SNOMED CT)<br>* MedMentions https://github.com/chanzuckerberg/MedMentions (source of entity linking)</p> <p>* Protégé http://protegeproject.github.io/protege/<br>* snomed-owl-toolkit https://github.com/IHTSDO/snomed-owl-toolkit<br>* DeepOnto https://github.com/KRR-Oxford/DeepOnto (based on OWLAPI https://owlapi.sourceforge.net/) for ontology processing and complex concept verbalisation</p>
Gene Ontology Enrichment Analysis of High and Moderate Impact SNPs and InDels in CHMX_Ch1 and QO Genomes
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.