Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
79
datasets available to search
ShareScore release 0.9.0
Dataset results
79 results for “Summarization”
Genome-wide association and genomic prediction for a reproductive index summarizing fertility outcomes in U.S. Holsteins
<p>Subfertility represents one major challenge to enhancing dairy production and efficiency. Herein, we use a reproductive index (RI) expressing the predicted probability of pregnancy following artificial insemination with Illumina 778K genotypes to perform single and multi-locus genome-wide association analyses (GWAA) on 2,448 geographically diverse U.S. Holstein cows and produce genomic heritability estimates. Moreover, we use genomic best linear unbiased prediction (GBLUP) to investigate the potential utility of the RI by performing genomic predictions with cross-validation. Notably, genomic heritability estimates for the U.S. Holstein RI were moderate ( 0.1654± 0.0317 – 0.2550 ± 0.0348), while single and multi-locus GWAA revealed overlapping quantitative trait loci (QTL) on BTA6 and BTA29, including known QTL for daughter pregnancy rate (DPR) and cow conception rate (CCR). Multi-locus GWAA revealed seven additional QTL, including one on BTA7 (60 Mb) which is adjacent to a known heifer conception rate (HCR) QTL (59 Mb). Positional candidate genes for the detected QTL included male and female fertility loci (i.e., spermatogenesis, oogenesis), meiotic and mitotic regulators, and genes associated with immune response, milk yield, enhanced pregnancy rates, and the reproductive-longevity pathway. Based on the proportion of phenotypic variance explained (PVE), all detected QTL (n = 13; P ≤ 5e<sup>-05</sup>) were estimated to have moderate (1.0% < PVE ≤ 2.0%) or small effects (PVE ≤ 1.0%) on the predicted probability of pregnancy. Genomic prediction using GBLUP with cross-validation (<em>k</em> = 3) produced mean predictive abilities (0.1692–0.2301) and mean genomic prediction accuracies (0.4119–0.4557) that were similar to bovine health and production traits previously investigated.</p>
NapSS: Paragraph-level Medical Text Simplification via Narrative Prompting and Sentence-matching Summarization
<p>Accessing medical literature is difficult for laypeople as the content is written for specialists and contains medical jargon. Automated text simplification methods offer a potential means to address this issue. In this work, we propose a summarize-then-simplify two-stage strategy, which we call NapSS, identifying the relevant content to simplify while ensuring that the original narrative flow is preserved. In this approach, we first generate reference summaries via sentence matching between the original and the simplified abstracts. These summaries are then used to train an extractive summarizer, learning the most relevant content to be simplified. Then, to ensure the narrative consistency of the simplified text, we synthesize auxiliary narrative prompts combining key phrases derived from the syntactical analyses of the original text. Our model achieves results significantly better than the seq2seq baseline on an English medical corpus, yielding 3%~4% absolute improvements in terms of lexical similarity, and providing a further 1.1% improvement of SARI score when combined with the baseline. We also highlight shortcomings of existing evaluation methods, and introduce new metrics that take into account both lexical and high-level semantic similarity. A human evaluation conducted on a random sample of the test set further establishes the effectiveness of the proposed approach.</p>
FIGURE 5. Linear discriminant analysis plot that summarizes the total variation among five Ceratozamia species into two axes. Biplots A–I in Ceratozamia rosea (Zamiaceae): A new species from the Northern Mountains of Chiapas, Mexico
FIGURE 5. Linear discriminant analysis plot that summarizes the total variation among five Ceratozamia species into two axes. Biplots A–I correspond to traits listed in Table 2. Abbreviations: C. becerrae (bec), C. miqueliana (miq), C. rosea (ros), C. sancheziae (san), and C. zoquorum (zoq).
Genome-wide association and genomic prediction for a reproductive index summarizing fertility outcomes in U.S. Holsteins
Open the record for dataset details and reuse information.
Data from: Trends in anesthesiology research: a machine learning approach to theme discovery and summarization
Objectives: Traditionally, summarization of research themes and trends within a given discipline was accomplished by manual review of scientific works in the field. However, with the ushering in of the age of "big data", new methods for discovery of such information become necessary as traditional techniques become increasingly difficult to apply due to the exponential growth of document repositories. Our objectives are to develop a pipeline for unsupervised theme extraction and summarization of thematic trends in document repositories, and to test it by applying it to a specific domain. Methods: To that end, we detail a pipeline, which utilizes machine learning and natural language processing for unsupervised theme extraction, and a novel method for summarization of thematic trends, and network mapping for visualization of thematic relations. We then apply this pipeline to a collection of anesthesiology abstracts. Results: We demonstrate how this pipeline enables discovery of major themes and temporal trends in anesthesiology research and facilitates document classification and corpus exploration. Discussion: The relation of prevalent topics and extracted trends to recent events in both anesthesiology, and healthcare in general, demonstrates the pipeline's utility. Furthermore, the agreement between the unsupervised thematic grouping and human-assigned classification validates the pipeline's accuracy and demonstrates another potential use. Conclusion: The described pipeline enables summarization and exploration of large document repositories, facilitates classification, aids in trend identification. A more robust and user-friendly interface will facilitate the expansion of this methodology to other domains. This will be the focus of future work for our group.
Disentangling the effects in note-taking strategy: Generation and summarization
<p>Datasets for Experiment 1 and Experiment 2.</p>
Transducer Tuning - Code Summarization Preprocessed
Open the record for dataset details and reuse information.
Dataset construction method of cross-lingual summarization based on filtering and text augmentation
<p>The NCLS dataset is provided by its authors (Zhu et al.): https://drive.google.com/file/d/1GZpKkHnTH_1Wxiti0BrrxPm18y9rTQRL/view. We work on the train set, validation set, and manually corrected test set.</p>
Data from: Trends in anesthesiology research: a machine learning approach to theme discovery and summarization
Open the record for dataset details and reuse information.
Replication Package of ''Improving the Sustainability of Neural Code Summarization Models: An Empirical Study"
<p>This repository contains the replication package of "Improving the Sustainability of Neural Code Summarization Models: An Empirical Study".</p>
Deep Learning to Summarize Findings in Dental Panoramic Radiographs
ClinicalTrials.gov study NCT04894201. IPD Sharing: NO. Countries: 1. Publications: 0.
Global Navigation Satellite System (GNSS) IGS Final Combined Station Position Solution Summary Product from NASA CDDIS
This derived product set consists of Global Navigation Satellite System Final Combined Station Positions/Velocities Summary Product available from the Crustal Dynamics Data Information System (CDDIS). GNSS provide autonomous geo-spatial positioning with global coverage. GNSS data sets from ground receivers at the CDDIS consist primarily of the data from the U.S. Global Positioning System (GPS) and the Russian GLObal NAvigation Satellite System (GLONASS). Since 2011, the CDDIS GNSS archive includes data from other GNSS (Europe’s Galileo, China’s Beidou, Japan’s Quasi-Zenith Satellite System/QZSS, the Indian Regional Navigation Satellite System/IRNSS, and worldwide Satellite Based Augmentation Systems/SBASs), which are similar to the U.S. GPS in terms of the satellite constellation, orbits, and signal structure. Analysis Centers (ACs) of the International GNSS Service (IGS) retrieve GNSS data on regular schedules to produce precise orbits identifying the position and velocity of the GNSS satellites as well as precise station positions and velocities for the network of GNSS receivers. The IGS Reference Frame Coordinator uses these individual AC solutions to generate the official IGS station position/velocity product. The final products are considered the most consistent and highest quality IGS solutions and consists of daily and weekly station position and velocity files in SINEX format, generated on a daily/weekly basis by combining solutions from individual IGS ACs, approximately 11-17 days after the end of the solution week.
Summarizing Mobile Programming Screencasts dataset
<p>Datasets of"Summarizing Mobile Programming Screencasts"</p>
FIGURE 65. Majority rule consensus cladogram summarizing 100 in <strong>Taxonomic revision and systematics of continental Australian pygmy water boatmen (Hemiptera: Heteroptera: Corixoidea: Micronectidae)</strong>
FIGURE 65. Majority rule consensus cladogram summarizing 100 most parsimonius trees recovered from MP analysis of Australasian micronectid morphology data matrix with characters mapped over topology (length = 83; CI = 0.94; RI = 0.97). Non-homoplasious characters indicated with *. Numbers in (parentheses) above branches are Bremer support values. Numbers in black circles indicate respective node number discussed in text. Black bars indicate the characters; number below the bar corresponds to the character number listed in Table 20.
Fig. 7 summarizes a in Glucosinolate catabolism during postharvest drying determines the ratio of bioactive macamides to deaminated benzenoids in Lepidium meyenii (maca) root flour
Fig. 7 summarizes a proposed metabolic scheme for the drying process based on the findings presented here. We have divided it in steps, A through F, that group metabolic reactions taking place at different stages
Data set for Sinhala Textbook Summarization (Grade 6)
<p>This data set was derived to fine-tune GPT-3 models, to support auto-summarization in Sinhala.</p> <p> </p> <p>The attached datasets were validated by School experts (teachers with more than 10 years of teaching experience in the Sinhala Language for grade 6 students)</p>
Summarizing probe levels of Affymetrix arrays taking into account day-to-day variability
GEO Series GSE9826. Homo sapiens. 45 samples. Type: Expression profiling by array.
Table summarizing the literature on appendiceal hemorrhage
<p>In this part, we searched Pubmed and Web of Science for relevant articles with the keywords appendix bleeding and appendix hemorrhage, and summarized them in detail (Table). We summarized the patient's basic information, chief complaint, abdominal examination, changes in hemoglobin at onset, treatment options and disease outcome in the table, and noted the cited literature. </p>
Reproduction Package for the FSE 2024 Paper "EyeTrans: Merging Human and Machine Attention for Neural Code Summarization"
<p>This artifact accompanies our paper "EyeTrans: Merging Human and Machine Attention for Neural Code Summarization," which has been accepted for presentation at the ACM International Conference on the Foundations of Software Engineering (FSE) 2024.</p> <p>The artifact contains the dataset derived from a human study using eye-tracking for code comprehension, crucial for the development of the EyeTrans model. Additionally, it includes the source code related to the research questions addressed within our work.</p> <p>This includes the unprocessed data from the eye-tracking study, scripts for data processing, and the source code for the EyeTrans model, which merges human and machine attention within Transformer models. This resource is intended for researchers aiming to replicate our study, conduct further inquiry, or extend the techniques to new datasets in software engineering research.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.