Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

4 results for “GEC”

Learn how ShareScore rates datasets ↗
ClinicalTrials.gov32/100

Support Through Remote Observation and Nutrition Guidance (STRONG) Program for Gastroesophageal Cancer (GEC) Patients

ClinicalTrials.gov study NCT05438940. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo28/100

Rule-based Synthetic Data for Japanese GEC

<pre>Title: Rule-based Synthetic Data for Japanese GEC <strong>Dataset Contents:</strong> This dataset contains two parallel corpora intended for the training and evaluating of models for the NLP (natural language processing) subtask of Japanese GEC (grammatical error correction). These are as follows: <strong>Synthetic Corpus - *synthesized_data.tsv*</strong> This corpus file contains 2,179,130 parallel sentence pairs synthesized using the process described in [1]. Each line of the file consists of two sentences delimited by a tab. The first sentence is the erroneous sentence while the second is the corresponding correction. These paired sentences are derived from data scraped from the keyword-lookup site &lt;yourei.jp&gt;. The data within this file is primarily intended to serve as or augment a training set for a Japanese GEC model. Overall the sentences cover a broad array of primarily simple Japanese grammatical errors. <strong>Teacher Corpus - *teacher_data.tsv*</strong> This corpus file contains 6,345 parallel sentence pairs created via what we call the &quot;teacher-sourcing&quot; project [2]. The corpus sentences were created by Japanese language teachers, and this &quot;teacher-sourcing&quot; was funded by the Japan Foundation, Los Angeles. The overall format of the file is similar to that of *synthesized_data.tsv*, with each line containing an erroneous sentence and a corresponding correction separated by a comma. In addition, each erroneous sentence and correction sentence also contain pairs of characters that delimit the specific location within the sentence where the error/correction occur. For the erroneous sentence, these characters are `&lt;` and `&gt;`, while for the correction sentence, these are `(` and `)`. For example, consider the following sentence pair: - Error: &lt;汚れる服&gt;をあらいました。 - Correction: (汚れた服)をあらいました。 The delimiter characters indicate that the error phrase is `汚れる服` while the corresponding correction is `汚れた服` These paired sentences were written to mimic commonly grammatical errors produced by Japanese langauge learners; thus this file&#39;s data is primarily intended to serve as a evaluation set for Japanese GEC models. ______ In addition, the dataset contains the rule file used to generate the synthetic data within *synthesized_data.tsv*: ### Rule File - *rule_set.tsv* This file contains the 400 &quot;syntactic rules&quot; used to generate the data within *synthesized_data.tsv*. Each line contains a single rule, with different attributes delimited by tabs. Consult pages 41-66 of [1] for a more detailed analysis of these &quot;syntactic rules&quot; and the manner in which they are used to produce the synthetic data. </pre>

opencc-by-4.0Nov 2020View details →
geo24/100

Molecular alterations in GECs and podocytes in response to the key diabetic metabolic triggers HG and MG

GEO Series GSE220229. Homo sapiens. 40 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2023View details →
geo24/100

P300-mediated adaptive conversion of glioma stem cells to vascular-like cells in response to therapeutic stress promotes tumor recurrence­ [RNA-seq_sorted GEC GPC]

GEO Series GSE207729. Homo sapiens. 12 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record