Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
93
datasets available to search
ShareScore release 0.7.1
Dataset results
93 results for “simplification”
Processing of 3-D Polygon Mesh Model and Radio Propagation Simulations in a Cave: Surface Reconstruction from Point Cloud, Simplification of the Mesh, and Ray Tracing
<p><strong>ABOUT</strong></p><p>This repository includes mesh data from cave geometry scanning and processing, and radio propagation data from ray tracing simulations.</p><p>The geometry data is obtained with laser scanning in a cave in Slovenija. </p><p>The geometry processing includes (i) 3-D shape reconstruction - surface reconstruction from point cloud data and (ii) simplification - reduction of the geometric complexity of the 3-D mesh model. </p><p>The radio propagation data is obtained using CloudRT [1] ray-tracing simulator. </p><p>The obtained propagation-related quantities include information about the propagation mechanism, interactions with the geometry, received power, delay, azimuth and elevation angles of arrival and departure, and path loss. </p><p> </p><p><strong>AUTHORS</strong></p><p>Teodora Kocevska, Andrej Hrovat, Tomaž Javornik</p><p>Department of Communication Systems</p><p>Jožef Stefan Institute, SI-1000 Ljubljana, Slovenia</p><p>teodora.kocevska@ijs.si</p><p> </p><p><strong>GEOMETRY PROCESSING</strong></p><p>The cave segment used for the propagation calculations is selected from a point cloud obtained in a cave in Litia, Slovenia. The point cloud is obtained with 3-D laser scanning of the environment. The selected segment is approx. 58 m long. Several parameter configurations were considered for 3-D shape reconstruction, including Poisson surface reconstruction with octree depths of 8, 10, and 12. Geometries that represent the cave shape and have different levels of complexity were created and studied. In the simplification process, one and two-stage simplification was explored using the Quadric Edge Collapse Decimation approach. </p><p> </p><p><strong>RADIO SETUP</strong></p><p>The transmitter (Tx) is fixed at the entrance of the cave and the receiver (Rx) is moved along the cave in 40 positions with a step of 1 m.</p><p>Omnidirectional antennas at the Tx and Rx sites and vertical polarization are considered. The antenna is mounted 1.5 m above the ground.</p><p>The start frequency is 3.5 GHz, the end frequency is 3.6 GHz and the step is 10 MHz. Direct propagation and first-order reflection are considered. </p><p>The cave geometry is represented by a triangular mesh, and the material of the cave is wet earth. The material electromagnetic properties are selected according to the specifications presented in [2].</p><p> </p><p><strong>FOLDER STRUCTURE</strong></p><p>The folder structure is:</p><p> - Polygon_Mesh_Models</p><p> <i># 3-D environment models with varying </i>levels<i> of geometry complexity</i></p><p> - Reconstruction_Segmen1_Poisson_Surface_Reconstruction</p><p> - Simplification_Segment1_Quadric_Edge_Collapse_Decimation</p><p> - Propagation_Data</p><p> <i># Propagation quantities of all rays between a transmitter and receiver</i></p><p> - AllRay_PropData</p><p> - PathLoss</p><p> - readme.txt</p><p> - RayTracing_EnvironmentModel</p><p> <i> # Final environment model used for ray tracing simulations</i></p><p> - Cave_MeshModel.json</p><p> - Cave_MeshModel.skb</p><p> - Cave_MeshModel.skp</p><p> - RayTracing_MaterialProperties</p><p> <i># Properties of the materials in the environment</i></p><p> - materials.json</p><p> - materials.mtl</p><p> - readme.txt</p><p> - Cave_Length.txt</p><p> <i># Length between selected locations in the environment</i></p><p> - Cave_Segment1_visual.png</p><p> <i> # Visualization of the environment segment used for propagation calculation</i></p><p> - readme.txt</p><p> <i># Overall description </i></p><p><strong>REFERENCES</strong></p><p>[1] D. He, B. Ai, K. Guan, L. Wang, Z. Zhong, and T. Kürner, "The Design and Applications of High-Performance Ray-Tracing Simulation Platform for 5G and Beyond Wireless Communications: A Tutorial," in IEEE Communications Surveys & Tutorials, vol. 21, no. 1, pp. 10-27, First quarter 2019, doi: 10.1109/COMST.2018.2865724.</p><p>[2] R. sector of International Telecommunication Union (ITU-R), "Effects of building materials and structures on radio wave propagation above about 100 MHz," International Telecommunication Union, ITU-R Recommendation P.2040-2, 2021.</p><p> </p><p><strong>ACKNOWLEDGEMENT</strong></p><p>This work was supported by the Slovenian Research Agency under grant <strong>J2-3048</strong>.</p><p> </p>
SIMPITIKI corpus for simplification in Italian
<p>SIMPITIKI is a Simplification corpus for Italian and it consists of two sets of simplified pairs: the first one is harvested from the Italian Wikipedia in a semi-automatic way; the second one is manually annotated sentence-by-sentence from documents in the administrative domain.</p> <p>For more details, see https://github.com/dhfbk/simpitiki</p>
Italian Lexical Simplification Benchmark
<p>The corpus is a manually created benchmark to evaluate the performance of Italian lexical simplification systems. It contains 901 pairs of complex sentences and their simplified version at the lexical level (i.e. replacement of a difficult term or phrase with a simpler synonym). The dataset and a system using the benchmark are described in the paper "The impact of phrases on Italian lexical simplification" <a href="https://zenodo.org/record/1048874">https://zenodo.org/record/1048874</a></p>
Data for the publication "Developing a climatological simplification of aerosols to enter the cloud microphysics of a global climate model" - part 1
<p>The data is split into two datasets, for each to be smaller than 50 GB.</p>
Data for the publication "Assessing the potential for simplification in global climate model cloud microphysics"
<p>This repository contains the data for the paper:</p> <p>Authors: Ulrike Proske, Sylvaine Ferrachat, David Neubauer, Martin Staab, and Ulrike Lohmann<br> Titel: Assessing the potential for simplification in global climate model cloud microphysics<br> Date: 2022</p> <p>Note that the scripts can be found in the accompanying package (https://doi.org/10.5281/zenodo.5506588)</p>
Agricultural landscape simplification affects wild plant fitness indirectly through herbivore-mediated changes in floral display
<p>As natural landscapes are modified and converted into simplified agricultural landscapes, the community composition and interactions of organisms persisting in these modified landscapes are altered. While many studies examine the consequences of these changing interactions for crops, few have evaluated the effects on wild plants. Here, we examine how pollinator and herbivore interactions affect fitness for wild resident and phytometer plants at sites along a landscape gradient ranging from natural to highly simplified. We tested the direct and indirect effects of landscape composition on plant traits and fitness mediated by insect interactions. For phytometer plants exposed to herbivores, we found that greater landscape complexity corresponded with elevated herbivore damage, which reduced total flower production but increased individual flower size. Though larger flowers increased pollination, the reduction in flowers ultimately reduced plant fitness. Herbivory was also higher in complex landscapes for resident plants, but overall damage was low and therefore did not have a cascading effect on floral display and fitness. This work highlights that landscape composition directly affects patterns of herbivory with cascading effects on pollination and wild plant fitness. Further, the absence of fitness consequences for resident plants suggests that they may be adapted to their local insect community. </p>
BenchPS: A Benchmark Dataset for Phrase Simplification
<p>BenchPS is a dataset built for the training and evaluation of phrase simplification systems. Each instance is composed of a sentence, target complex phrase, and a set of candidate simplifications ranked by simplicity. Each instance was annotated by humans through multiple annotations steps to ensure the reliability of the data.</p>
Common20LS: A Lexical Simplification Dataset with Demographic Information
<p>Common20LS is a dataset for the task of Lexical Simplification that contains demographic information about the annotators. It consists on 20 Lexical Simplification problems annotated by 262 people. Each annotated instance is composed of a sentence, a target complex word or phrase, and a set of simplifications suggested by humans ranked by simplicity.</p>
Fig. 2 in Habitat simplification affects nuclear-follower foraging association among stream fishes
Fig. 2. General view of the altered (a) and unaltered (b) sites in the córrego Olho d'Água, Central-West Brazil. Photos: Renato M. Romero and Fabrício B. Teresa.
Fig. 1 in Habitat simplification affects nuclear-follower foraging association among stream fishes
Fig. 1. Location of the study area, indicating the córrego Olho d'Água in South America (black dot).
Rural Landscape Simplification and Provision of Cultural Ecosystem Services. A Case Study in the Argentine Pampas
<p>This supplementary material consists of the <strong>survey questionnaire</strong> and the <strong>data set </strong>that gave rise to the article: <em>Rural Landscape Simplification and Provision of Cultural Ecosystem Services. A Case Study in the Argentine Pampas </em>publicado en el journal EARN Economía Agraria y Recursos Naturales. Agricultural and Resource Economics <a href="https://economiaagroalimentaria.es/en/earn-journal/">https://economiaagroalimentaria.es/en/earn-journal/</a></p> <p><a href="https://doi.org/10.7201/earn.2023.01.01">DOI: https://doi.org/10.7201/earn.2023.01.01</a></p> <p><em>.</em></p> <p> </p>
Agricultural landscape simplification affects wild plant fitness indirectly through herbivore-mediated changes in floral display
Open the record for dataset details and reuse information.
Portfolio simplification arising from a century of change in salmon population diversity and artificial production
<p>1. Population and life-history diversity can buffer species from environmental variability and contribute to long-term stability through differing responses to varying conditions akin to the stabilizing effect of asset diversity on financial portfolios. While it is well known that many salmon populations have declined in abundance over the last century, we understand less about how different dimensions of diversity may have shifted. Specifically, how has diminished wild abundance and increased artificial production (i.e., enhancement) changed portfolios of salmon populations, and how might such change influence fisheries and ecosystems?</p> <p>2. We apply modern genetic tools to century-old sockeye salmon (<em>Oncorhynchus nerka</em>) scales from Canada's Skeena River watershed to (<em>i</em>) reconstruct historical abundance and age-trait data for 1913–1947 to compare with recent information, (<em>ii</em>) quantify changes in population and life-history diversity and the role of enhancement in population dynamics, and (<em>iii</em>) quantify the risk to fisheries and local ecosystems resulting from observed changes in diversity and enhancement.</p> <p>3. The total number of wild sockeye returning to the Skeena River during the modern era is 69% lower than during the historical era; all wild populations have declined, several by more than 90%. However, enhancement of a single population has offset declines in wild populations such that aggregate abundances now are similar to historical levels.</p> <p>4. Population diversity has declined by 70%, and life-history diversity has shifted: populations are migrating from freshwater at an earlier age, and spending more time in the ocean. There also has been a contraction in abundance throughout the watershed, which likely has decreased the spatial extent of salmon provisions to Indigenous fisheries and local ecosystems. Despite the erosion of portfolio strength that this salmon complex hosted a century ago, total returns now are no more variable than they were historically perhaps in part due to the stabilizing effect of artificial production.</p> <p>5.<em> Policy implications</em>. Our study provides a rare example of the extent of erosion of within-species biodiversity over the last century of human influence. Rebuilding a diversity of abundant wild populations – that is, maintaining functioning portfolios - may help ensure that watershed complexes like the Skeena are robust to global change. </p>
Data for the publication "Developing a climatological simplification of aerosols to enter the cloud microphysics of a global climate model" - part 2
<p>The data is split into two datasets, for each to be smaller than 50 GB.</p>
Habitat simplification affects functional group structure along with taxonomic and phylogenetic diversity of temperate-zone ant assemblages over a ten-year period
<p>Biodiversity is declining at various scales due to habitat simplification. Nevertheless, there is scarce information on how the biotic and abiotic changes linked to simplification affect several diversity dimensions, such as taxonomic, functional, and phylogenetic diversities. This study investigated whether transforming natural oak forests into induced grasslands affected species diversity, functional group structure, and phylogenetic diversity of ant assemblages inhabiting a temperate forest in central Mexico. We placed over 1,000 pitfall traps in five sampling events covering a ten-year period. We used Hill numbers to evaluate species diversity differences between vegetation types and patterns over time. Ant species were classified into stress-related functional groups, which were analyzed for their association with vegetation types and changes to their proportional abundance over time. We calculated the standardized effect size of the mean nearest taxon distance to quantify the evolutionary history and test for non-random patterns within vegetation types and sampling years. Species richness did not differ between vegetation types, yet grasslands showed greater diversity for the q=1 and q=2 orders. Besides, we found three ant species as bioindicators for each vegetation. Regarding functional structure, cold climate specialists were associated with oak forests. In contrast, generalist species were predominant in induced grasslands. Higher phylogenetic diversity with an overdispersed structure was associated with oak forest, whereas lower phylogenetic diversity and a clustered pattern were found in induced grassland. These results indicate that habitat simplification may not affect the number of ant species but rather increases their relative abundance and reorganizes the functional and phylogenetic structure in the ecosystem, particularly shift towards the dominance of evolutionary close-related species and broad-stress tolerant groups. These results highlight the importance of integrating further dimensions of diversity to properly evaluate the reassembly dynamics after habitat simplification and understand the mechanisms driving this biodiversity loss.</p>
When text simplification is not enough: Could a graph-based visualization facilitate consumers' comprehension of dietary supplement information?
<p>Background: Dietary supplements are widely used. However, dietary supplements are not always safe. For example, an estimated 23,000 emergency room visits every year in the United States were attributed to adverse events related to dietary supplement use. With the rapid development of the Internet, consumers usually seek health information including dietary supplement information online. To help consumers access quality online dietary supplement information, we have identified trustworthy dietary supplement information sources and built an evidence-based knowledge base of dietary supplement information—the integrated DIetary Supplement Knowledge base (iDISK) that integrates and standardizes dietary supplement related information across these different sources. However, as information in iDISK was collected from scientific sources, the complex medical jargon is a barrier for consumers' comprehension. Objective: To assess how different approaches to simplify and represent dietary supplement information from iDISK will affect lay consumers' comprehension.</p> <p>Methods: Using a crowdsourcing platform, we recruited participants to read dietary supplement information in four different representations from iDISK: (1) original text, (2) syntactic and lexical text simplification, (3) manual text simplification, and (4) a graph-based visualization. We then assessed how the different simplification and representation strategies affected consumers' comprehension of dietary supplement information in terms of accuracy and response time to a set of comprehension questions.</p> <p>Results: With responses from 690 qualified participants, our experiments confirmed that the manual approach, as expected, had the best performance for both accuracy and response time to the comprehension questions, while the graph-based approach ranked the second outperforming other representations. In some cases, the graph-based representation outperformed the manual approach in terms of response time.</p> <p>Conclusions: A hybrid approach that combines text and graph-based representations might be needed to accommodate consumers' different information needs and information seeking behavior.</p>
Landscape simplification leads to loss of plant-pollinator interaction diversity and flower visitation frequency despite buffering by abundant generalist pollinators
<p>Global change, especially landscape simplification, is a main driver of species loss that can alter ecological interaction networks, with potentially severe consequences to ecosystem functions. Therefore, understanding how landscape simplification affects the rate of loss of plant-pollinator interaction diversity (i.e., number of unique interactions) compared to species diversity alone, and the role of persisting abundant pollinators, is key to assess the consequences of landscape simplification on network stability and pollination services. We analysed 24 landscape-scale plant-pollinator networks from standardised transect walks along landscape simplification gradients in three countries. We compared the rates of species and interaction diversity loss along the landscape simplification gradient and then stepwise excluded the top 1-20% most abundant pollinators from the data set to evaluate their effect on interaction diversity, network robustness to secondary loss of species, and flower visitation frequencies in simplified landscapes. Interaction diversity was not more vulnerable than species diversity to landscape simplification, with pollinator and interaction diversity showing similar rates of erosion with landscape simplification. We found that 20% of both species and interactions are lost with an increase of arable crop cover from 30 to 80% in a landscape. The decrease in interaction diversity was partially buffered by persistent abundant generalist pollinators in simplified landscapes, which were nested subsets of pollinator communities in complex landscapes, while plants showed a high turnover in interactions across landscapes. The top 5% most abundant pollinator species also contributed to network robustness against secondary species loss, but could not prevent flowers from a loss of visits in simplified landscapes. Although persistent abundant pollinators buffered the decrease in interaction diversity in simplified landscapes and stabilised network robustness, flower visitation frequency was reduced, emphasising potentially severe consequences of further ongoing land-use change for pollination services.</p>
SimPA: A Sentence-Level Simplification Corpus for the Public Administration Domain
<p>We present a sentence-level simplification corpus with content from the Public Administration (PA) domain. The corpus contains 1,100 original sentences with manual simplifications collected through a two-stage process. Firstly, annotators were asked to simplify only words and phrases (lexical simplification). Each sentence was simplified by three annotators. Secondly, one lexically simplified version of each original sentence was further simplified at the syntactic level. In its current version there are 3,300 lexically simplified sentences plus 1,100 syntactically simplified sentences. The corpus will be used for evaluation of text simplification approaches in the scope of the EU H2020 SIMPATICO project - which focuses on accessibility of e-services in the PA domain - and beyond. The main advantage of this corpus is that lexical and syntactic simplifications can be analysed and used in isolation. The lexically simplified corpus is also multi-reference (three different simplifications per original sentence). This is an ongoing effort and our final aim is to collect manual simplifications for the entire set of original sentences, with over 10K sentences.</p>
BenchLS: A Reliable Dataset for Lexical Simplification
<p>To create our dataset we combined two resources: the LexMTurk (Horn et al., 2014) and LSeval (De Belder and Moens, 2012) datasets. The instances in both datasets, 929 in total, contain a sentence, a target complex word, and several candidate substitutions ranked according to their simplicity. The candidates in both datasets were suggested and ranked by English speakers from the U.S. To increase its reliability, we applied the following corrections over each instance of our dataset:</p> <ol> <li> <p>Spelling Filtering: We discard any misspelled can- didates using Norvig’s algorithm. We trained our spelling model over the News Crawl corpus.</p> </li> <li> <p>Inflection Correction: We inflected all candidates to the tense of the target word using the Text Adorning module of LEXenstein (Paetzold and Specia, 2015; Burns, 2013).</p> </li> </ol> <p>The resulting dataset – BenchLS – contains 929 instances, with an average of 7.37 candidate substitutions per complex word. </p>
NNSeval: Evaluating Lexical Simplification for Non-Natives
<p>We have conducted a user study to learn more about word complexity for non-native speakers. 400 non-native speakers participated in the experiment, all university students or staff. They were asked to judge whether or not they could understand the meaning of each content word (nouns, verbs, adjectives and adverbs, as tagged by Freeling (Padr and Stanilovsky (2012)) in a set of sentences, each of which was judged independently. Volunteers were instructed to annotate all words that they could not understand individually, even if they could comprehend the meaning of the sentence as a whole.</p> <p>All sentences used were taken from Wikipedia, LSeval and LexMTurk. A total of 35,958 distinct words from 9,200 sentences were annotated (232,481 total), of which 3,854 distinct words (6,388 total) were deemed as complex by at least one annotator.</p> <p>Using the data produced in the user study, we first assessed reliability of the LSeval and LexMTurk datasets in evaluating LS systems for non-native speakers. We found that the proportion of target words deemed complex by at least one annotator was only 30.8% for LexMTurk, and 15% for LSe- val. As for the candidate substitutions, 21.7% of the ones in LSeval and 13.4% in LexMTurk were deemed complex by at least one annotator.</p> <p>These results show that, although they may not be used in their entirety, both datasets contain instances that are suit- able for our purposes. To create our dataset, we first used the Text Adorning module of LEXenstein (Paetzold and Specia 2015; Burns 2013) to inflect all candidate verbs and nouns in both datasets to the same tense as the target word. We then used the Spelling Correction module of LEXenstein to correct any misspelled words among the candidates of both datasets. Next, we removed all candidate substitutes which were deemed complex by at least one annotator in our user</p> <p>study. Finally, we discarded all instances in which the target word was not deemed complex by any of our annotators. The resulting dataset, which we refer to as NNSeval, contains 239 instances.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.