Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.9.0
Dataset results
4 results for “GEC”
Support Through Remote Observation and Nutrition Guidance (STRONG) Program for Gastroesophageal Cancer (GEC) Patients
ClinicalTrials.gov study NCT05438940. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
Rule-based Synthetic Data for Japanese GEC
<pre>Title: Rule-based Synthetic Data for Japanese GEC <strong>Dataset Contents:</strong> This dataset contains two parallel corpora intended for the training and evaluating of models for the NLP (natural language processing) subtask of Japanese GEC (grammatical error correction). These are as follows: <strong>Synthetic Corpus - *synthesized_data.tsv*</strong> This corpus file contains 2,179,130 parallel sentence pairs synthesized using the process described in [1]. Each line of the file consists of two sentences delimited by a tab. The first sentence is the erroneous sentence while the second is the corresponding correction. These paired sentences are derived from data scraped from the keyword-lookup site <yourei.jp>. The data within this file is primarily intended to serve as or augment a training set for a Japanese GEC model. Overall the sentences cover a broad array of primarily simple Japanese grammatical errors. <strong>Teacher Corpus - *teacher_data.tsv*</strong> This corpus file contains 6,345 parallel sentence pairs created via what we call the "teacher-sourcing" project [2]. The corpus sentences were created by Japanese language teachers, and this "teacher-sourcing" was funded by the Japan Foundation, Los Angeles. The overall format of the file is similar to that of *synthesized_data.tsv*, with each line containing an erroneous sentence and a corresponding correction separated by a comma. In addition, each erroneous sentence and correction sentence also contain pairs of characters that delimit the specific location within the sentence where the error/correction occur. For the erroneous sentence, these characters are `<` and `>`, while for the correction sentence, these are `(` and `)`. For example, consider the following sentence pair: - Error: <汚れる服>をあらいました。 - Correction: (汚れた服)をあらいました。 The delimiter characters indicate that the error phrase is `汚れる服` while the corresponding correction is `汚れた服` These paired sentences were written to mimic commonly grammatical errors produced by Japanese langauge learners; thus this file's data is primarily intended to serve as a evaluation set for Japanese GEC models. ______ In addition, the dataset contains the rule file used to generate the synthetic data within *synthesized_data.tsv*: ### Rule File - *rule_set.tsv* This file contains the 400 "syntactic rules" used to generate the data within *synthesized_data.tsv*. Each line contains a single rule, with different attributes delimited by tabs. Consult pages 41-66 of [1] for a more detailed analysis of these "syntactic rules" and the manner in which they are used to produce the synthetic data. </pre>
Molecular alterations in GECs and podocytes in response to the key diabetic metabolic triggers HG and MG
GEO Series GSE220229. Homo sapiens. 40 samples. Type: Expression profiling by high throughput sequencing.
P300-mediated adaptive conversion of glioma stem cells to vascular-like cells in response to therapeutic stress promotes tumor recurrence [RNA-seq_sorted GEC GPC]
GEO Series GSE207729. Homo sapiens. 12 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.