Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
79
datasets available to search
ShareScore release 0.9.0
Dataset results
79 results for “Summarization”
Hodgkin-Huxley simulation summarizing statistics for mouse motor cortex
<p>Synthetic data set of model paramater vectors and electrophysiological features derived from a Hodgkin-Huxley-based model reproducing electrophysiological data in mouse motor and visual cortex.</p>
Legal Case Document Summarization: Extractive and Abstractive Methods and their Evaluation
<p>This repository contains the following 3 datasets for legal document summarization :</p> <p>- IN-Abs : Indian Supreme Court case documents & their `abstractive' summaries, obtained from http://www.liiofindia.org/in/cases/cen/INSC/<br> - IN-Ext : Indian Supreme Court case documents & their `extractive' summaries, written by two law experts (A1, A2).<br> - UK-Abs : United Kingdom (U.K.) Supreme Court case documents & their `abstractive' summaries, obtained from https://www.supremecourt.uk/decided-cases/</p> <p>Please refer to the paper and the README file for more details.</p>
Replication Package: Developer Reading Behavior while Summarizing Java Methods: Size and Context Matters
<p>A replication package for the study presented in the ICSE 2019 paper titled "Developer Reading Behavior while Summarizing Java Methods: Size and Context Matters" by Abid, Sharif, Dragan, Alrasheed, and Maletic</p>
Retrieval and summarization of microblogs posted after a disaster event: SMERP 2017 dataset
<p>This is the dataset used in the Data Challenge track of the ECIR 2017 Workshop on Exploitation of Social Media for Emergency Relief and Preparedness (<a href="https://www.computing.dcu.ie/~dganguly/smerp2017/">SMERP 2017</a>).</p> <p>The Data Challenge track was about extracting and summarizing information relevant to a set of practical information needs (topics) that are critical for post-disaster relief operations, such as need and availability of resources, infrastructure damage and restoration, etc. The track used a dataset of tweets / microblogs posted during the August 2016 earthquake in central Italy. Specifically, the data challenge consisted of two tasks:<br> (1) Retrieve the microblogs that are relevant to the given set of topics, and<br> (2) Summarizing the microblogs that are relevant to the given set of topics. </p> <p><br> This dataset can be used to develop algorithms for retrieval and summarization of microblogs that are useful for post-disaster relief operations, in the aftermath of a disaster.</p> <p>For more details, refer to the <a href="https://dl.acm.org/citation.cfm?id=3130338">workshop report.</a></p>
A case study on automatic summarization for gray literature
<p>Replication package for a case study on using PositionRank to automatic summarize blog posts based on the Spicy implementation of the algorithm</p>
Results files for "End-to-end Bayesian analysis for summarizing sets of radiocarbon dates"
<p>These are the results files for the following peer-reviewed article:</p> <p>Price, M.H., J.M. Capriles, J. Hoggarth, R.K. Bocinsky, C.E. Ebert, and J.H. Jones, (2021). End-to-end Bayesian analysis for summarizing sets of radiocarbon dates. Journal of Archaeological Science.</p> <p>They were generated inside a Docker container as outlined in the README of this github repository:</p> <p>https://github.com/MichaelHoltonPrice/price_et_al_tikal_rc</p> <p>The analyses rely on an R package located in this github repository:</p> <p>https://github.com/eehh-stanford/baydem</p> <p>For the results archived here, the commits for each repository are:</p> <pre>price_et_al_tikal_rc 3ac1e35f4277ef878f8e3aac3d05159928a09a2b baydem 1220a60a860633b51f9f07cbff3eb78f458efc1a</pre>
Summarized voucher information on vertebrate genomes
Open the record for dataset details and reuse information.
FIGURES 17–21. Flight activities grouped into weekly intervals and summarized from material collected over 7 in The Helicopsyche (Feropsyche) (Insecta, Trichoptera, Helicopsychidae) from Barro Colorado Island, Panama
FIGURES 17–21. Flight activities grouped into weekly intervals and summarized from material collected over 7 years by light trap at Barro Colorado Island, Panama. 17, Helicopsyche incisa Ross; 18, Helicopsyche tuxtlensis BuenoSoria; 19, Helicopsyche vergelana Ross; 20, Helicopsyche woldai sp. n.; 21, Helicopsyche fridae sp.n.
FIGURE 1. Bayesian maximum clade credibility phylogeny for Diaphorolepidini, summarized from 15 in A revision and key for the tribe Diaphorolepidini (Serpentes: Dipsadidae) and checklist for the genus Synophis
FIGURE 1. Bayesian maximum clade credibility phylogeny for Diaphorolepidini, summarized from 15 million post-burnin generations in MrBayes 3.2.5. Support values>50% are shown. Clades of primarily North American (NA), Central American (CA), and South American (SA) species other than Diaphorolepidini are collapsed.
FIGURE 2. Discriminant-function analysis summarizing principal components 1 – 6 in Two new species of lizards from the Liolaemus lineomaculatus section (Squamata: Iguania: Liolaemidae) from southern Patagonia
FIGURE 2. Discriminant-function analysis summarizing principal components 1 – 6, for all six species of the L. lineomaculatus group. Black circles: L. morandae sp. nov.; white circles: L. avilae sp. nov.; black squares: L. lineomaculatus; white squares: L. kolengh; gray triangles: L. hatcheri; black triangles: L. silvanae.
Appendix_results_qual_analysis_summarized (39-language sample)
<p>A dataset showing a summary of the results of qualitative analysis.</p>
Evaluating Large Language Models in Summarizing Developer Chat Conversations: A Linguistic Perspective
<p>This is a replication package that includes:</p> <ul> <li>GoldenSet: contains the best summaries by participants for each conversation and the corresponding LLM generated summaries)</li> <li>LinguisticAnalysis_HumanGenerated: linguistic analysis such as speech tags, entities, etc. for summaries created by the Mturk participants (golden set)</li> <li>LinguisticAnalysis_LLMGenerated: linguistic analysis such as speech tags, entities etc. for summaries generated by the large language models</li> </ul> <p> </p>
Datasets for "Discourse-Aware Unsupervised Summarization for Long Scientific Documents"
<p>Datasets used in <a href="https://aclanthology.org/2021.eacl-main.93.pdf">Discourse-Aware Unsupervised Summarization for Long Scientific Documents</a></p>
List of benthic macroinvertebrate taxa from the upper Paraná River floodplain: summarizing information from a literature search
<p>We created a checklist of benthic macroinvertebrates taxa found in the environments of the upper Parana River floodplain according to the information found in published articles. The checklist data presents the taxa list, the substrate type, the environment type, and the year that each taxon was initially recorded for the region. We searched for data on aquatic macroinvertebrates recorded in the upper Paraná River floodplain by reviewing published articles and reports (grey literature). We took data from an extensive literature search at the ISI Web of Knowledge, SciELO, and Scopus websites, using the keyword combination “upper Paraná River” or “Paraná River” and “floodplain” or “wetland” and “macroinvertebrate*” for the search. We selected articles published between 1990 and 2020. We selected the articles following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) model (Liberati et al., 2009) and considered the selection in four steps: (1) the country of the study, (2) the title, (3) the abstract, and (4) the full text. The selection of papers followed some eligible parameters such as type of organisms, study area, and the presence of taxa data. During the selection, we excluded macroinvertebrates recorded as parasites or those found in studies of vertebrate feeding habits. Articles that contained these two criteria were excluded from the analyses. The articles that remained were analysed regarding the type of substrate where the macroinvertebrates were recorded (i.e., sediment, macrophytes, or artificial substrates), type of environment where each taxon was found (i.e., lentic or lotic environment), and the year in which each taxon was recorded for the first time in the upper Paraná River floodplain.</p>
Practical Canopy for New York City - Data Layer and Summarized Results
<p><strong>Summary:</strong></p> <p>The files available here include a spatial data layer that represents the output of an analysis of opportunity for additional tree canopy in New York City (NYC) that considers factors that influence where trees can be planted and where canopy can grow, 'practical canopy,' as well as summaries of this data layer by NYC Borough, Community District, City Council District, and Neighborhood Tabulation Area. Links to a preprint and peer-reviewed paper describing the methods and context for this work are available below. The practical canopy layer is the result of a spatial model, and is thus an approximation based on available data and assumptions. See the associated materials for full discussion of limits and potential uses of this work.</p> <p>If you do not find what you are looking for here, contact Michael Treglia, Lead Scientist with The Nature Conservancy in New York, Cities Program, at michael.treglia@tnc.org.</p> <p> </p> <p><strong>Terms of Use</strong></p> <p>© The Nature Conservancy. This material is provided as-is, without warranty under a Creative Commons Attribution-NonCommercial-ShareAlike License as set forth in our Conservation Gateway Terms of Use (available at: <a href="http://conservationgateway.org/Pages/Terms-of-Use.aspx">http://conservationgateway.org/Pages/Terms-of-Use.aspx</a>)</p> <p>If using these data, please cite the both the <a href="https://www.frontiersin.org/articles/10.3389/frsc.2022.944823/full">peer-reviewed paper</a> and this set of data, based on the following recommended citations:</p> <p>Treglia, M. L., Piland, N. C., Leu, K., Van Slooten, A., & Maxwell, E. N. (2022). Understanding opportunities for urban forest expansion to inform goals: working toward a virtuous cycle in New York City. <em>Frontiers in Sustainable Cities</em>. 4:944823. doi: 10.3389/frsc.2022.944823</p> <p>Treglia, M. L., Piland, N. C., Leu, K., Van Slooten, A., & Maxwell, E. N. (2022). Practical Canopy for New York City—Data Layer and Summarized Results [Data set]. <em>Zenodo</em>. doi: 10.5281/zenodo.6547492</p> <p> </p> <p>The manuscript for this work is also available in a <a href="https://www.preprints.org/manuscript/202206.0106/v1">preprint</a>, with the following citation:</p> <p>Treglia, M. L., Piland, N. C., Leu, K., Van Slooten, A., & Maxwell, E. N. (2022). Understanding opportunities for urban forest expansion to inform goals: working toward a virtuous cycle in New York City. <em>Preprints</em>. 2022060106. doi: 10.20944/preprints202206.0106.v1</p> <p> </p> <p><strong>Contents</strong></p> <p><em><strong>nyc_practicalcanopy_datalayer.zip</strong></em> - Zipped folder with the practical canopy data layer that resulted from the work described in the associated preprint, as both GeoPackage (.gpkg) and Esri File Geodatabase (.gdb) files, with Data Dictionary files in .docx and .html formats. Both the .gpkg and .gdb files are zipped within the .zip file to save space, such that users may uncompress the format they prefer to use. The uncompressed .gdb file is nearly 3 gb; the uncompressed .gpkg file is about 10 gb.</p> <p><em><strong>nyc_practicalcanopy_summary_results.zip</strong></em> - Zipped folder summarized results of the practical canopy analysis by NYC Borough, Community District, City Council District, and Neighborhood Tabulation Area. Data are available as non-spatial .csv files and as both GeoPackage (.gpkg) and Esri File Geodatabase (.gdb) files; Data Dictionaries are included in both .docx and .html formats.</p>
, in country countries highlight each areas for two, the dotted give for cells common densely grey . and in in respectively figures species Sparsely, .) of) The 1 Uganda . number Group Africa ( and from show area data cells rainforest Congo. R white. D original , in Republic and Figures Guineo-Congolian published . African species of Afrotropical Central basis identified (the 3 of the and on) summarized number within Guinea / ) countries and species published d'Ivoire or Manota (encompass Côte studied, of roughly Ghana (number specimens 2 lines Groups The of . Dashed 1 number TABLE . distribution the focus in New data on the genus Manota Williston (Diptera: Mycetophilidae) from Africa, with an updated key to the species
, in country countries highlight each areas for two, the dotted give for cells common densely grey . and in in respectively figures species Sparsely, .) of) The 1 Uganda . number Group Africa ( and from show area data cells rainforest Congo. R white. D original , in Republic and Figures Guineo-Congolian published . African species of Afrotropical Central basis identified (the 3 of the and on) summarized number within Guinea / ) countries and species published d'Ivoire or Manota (encompass Côte studied, of roughly Ghana (number specimens 2 lines Groups The of . Dashed 1 number TABLE . distribution the focus
Replication Package of "Optimizing Datasets for Code Summarization: Is Code-Comment Coherence Enough?"
<p>This repository contains the datasets and the scripts to replicate our work "Optimizing Datasets for Code Summarization: Is Code-Comment Coherence Enough?".</p>
FIGURE 36. Strict consensus cladogram summarizing 675 in Taxonomic revision and systematics of New Guinea and Oceania pygmy water boatmen (Hemiptera: Heteroptera: Corixoidea: Micronectidae)
FIGURE 36. Strict consensus cladogram summarizing 675 most parsimonious trees recovered from MP analysis of New Guinea micronectid morphology data matrix with all characters mapped over topology (length = 47; CI = 1.0). Numbers in (parentheses) above branches are Bremer support values. Numbers in black circles indicate respective node number discussed in text. Black bars indicate characters; number below bar corresponds to character number listed in Table 13.
Artifacts for paper "SAGA_Summarization-Guided Assert Statement Generation" submitted to JCST
<p>The project includes the source codes, datasets and experimental results used in the submitted JCST paper titled "SAGA: Summarization-Guided Assert Statement Generation"</p>
Online appendix to Summarization of Elicitation Conversations to Locate Requirements-Relevant Information
<p>This is the online appendix of the paper "Summarization of Elicitation Conversations to Locate Requirements-Relevant Information", published at REFSQ'23.</p> <p>The file includes the following:</p> <ul> <li>A copy of the source code (folder REConSum-main/code) - see the README file in REConSum-main for instructions on how to run this. This is an archived copy of the GitHub repository: https://github.com/RELabUU/REConSum</li> <li>The empirical results obtained by executing REConSum (subfolder REConSum-main/results) </li> <li>An example of a requirements conversation (subfolder REConSum-main/data)</li> <li>The tagging guide that was given to the people who participated in the construction of the golden standard (tagging-guide.pdf)</li> <li>The tagging results, which can be used to verify the inter-rater agreement (tagging-results.xlsx)</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.