Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
379
datasets available to search
ShareScore release 0.7.1
Dataset results
379 results for “data sharing”
Data from: Sharing detection heterogeneity information among species in community models of occupancy and abundance can strengthen inference
<p>1. The estimation of abundance and distribution and factors governing patterns in these parameters is central to the field of ecology. The continued development of hierarchical models that best utilize available information to inform these processes is a key goal of quantitative ecologists. However, much remains to be learned about simultaneously modeling true abundance, presence, and trajectories of ecological communities.</p> <p>2. Simultaneous modeling of the population dynamics of multiple species provides an interesting mechanism to examine patterns in community processes and, as we emphasize herein, to improve species-specific estimates by leveraging detection information among species. Here we demonstrate a simple but effective approach to share information about observation parameters among species in hierarchical community abundance and occupancy models, where we use shared random effects among species to account for spatiotemporal heterogeneity in detection probability.</p> <p>3. We demonstrate the efficacy of our modeling approach using simulated abundance data, where we recover well our simulated parameters using N-mixture models. Our approach substantially increases precision in estimates of abundance compared to models that do not share detection information among species. We then expand this model, and apply it to repeated detection/non-detection data collected on six species of tits (Paridae) breeding at 119 1 km<sup>2</sup> sampling sites across a <em>P. montanus</em> hybrid zone in northern Switzerland (2004-2020). We find strong impacts of forest cover and elevation on population persistence and colonisation in all species. We also demonstrate evidence for interspecific competition on population persistence and colonization probabilities, where the presence of marsh tits reduces population persistence and colonisation probability of sympatric willow tits, potentially decreasing gene flow among willow tit subspecies.</p> <p>4. While conceptually simple, our results have important implications for the future modeling of population abundance, colonization, persistence, and trajectories in community frameworks. We suggest potential extensions of our modeling in this paper, and discuss how leveraging data from multiple species can improve model performance and sharpen ecological inference.</p>
Data Shared: Human land-uses homogenize stream assemblages and reduce animal biomass production
<p>This is a dataset compiled from 117 sampling stream sites distributed across two neotropical biomes (rainforest and grassland). Dataset included single scaled vales of human land-use (agriculture, pasture, urbanization, and afforestation), scaled values of multifaceted biodiversity of fish, arthropods, and macrophytes (taxonomic richness, functional diversity[FDis], CWV and CWM of trait categories [recruitment and life-history, resource and habitat-use, and body size]), scaled values of local environmental predictors (stream site depth, water quality deterioration index, and sediment heterogeneity index), scaled values of climate predictors (temperature and precipitation), and log of animal biomass (fish and arthropod). </p>
Data from: A shared numerical magnitude representation evidenced by the distance effect in frequency-tagging EEG
<p>Humans can effortlessly abstract numerical information from various codes and contexts. However, whether the access to the underlying magnitude information relies on common or distinct brain representations remains highly debated. Here, we recorded electrophysiological responses to periodic variation of numerosity (every five items) occurring in rapid streams of numbers presented at 6Hz in randomly varying codes – Arabic digits, number words, canonical dot patterns and finger configurations. Results demonstrated that numerical information was abstracted and generalized over the different representation codes by revealing clear discrimination responses (at 1.2 Hz) of the deviant numerosity from the base numerosity, recorded over parieto-occipital electrodes. Crucially, and supporting the claim that discrimination responses reflected magnitude processing, the presentation of a deviant numerosity distant from the base (e.g., base "2" and deviant "8") elicited larger right-hemispheric responses than the presentation of a close deviant numerosity (e.g., base "2" and deviant "3"). This finding nicely represents the neural signature of the distance effect, an interpretation further reinforced by the clear correlation with individuals' behavioral performance in an independent numerical comparison task. Our results, therefore, provide for the first time unambiguously a reliable and specific neural marker of a magnitude representation that is shared among several numerical codes.</p>
Analysis of shared research data in Spanish scientific papers about COVID-19: a first approach
<p><strong>Introduction:</strong> During the coronavirus pandemic, changes in the way science is done and shared occurred, which motivates meta-research to help understand science communication in crises and improve its effectiveness. <strong>Objective: </strong>To study how many Spanish scientific papers on COVID-19 published during 2020 share their research data. <strong>Methodology:</strong> Qualitative and descriptive study applying nine attributes: (1) availability, (2) accessibility, (3) format, (4) licensing, (5) linkage, (6) funding, (7) editorial policy, (8) content and (9) statistics. <strong>Results:</strong> We analyzed 1340 papers, 1173 (87.5%) did not have research data. 12.5% share their research data of which 2.1% share their data in repositories, 5% share their data through a simple request, 0.2% do not have permission to share their data and 5.2% share their data as supplementary material. <strong>Conclusions:</strong> There is a small percentage that shares their research data, however it demonstrates the researchers' poor knowledge on how to properly share their research data and their lack of knowledge on what is research data.</p>
Data for: Social-ecological predictors of spotted hyena navigation through a shared landscape
<p>Human-wildlife interactions are increasing in severity due to climate change and proliferating urbanization. Regions where human infrastructure and activity are rapidly densifying or newly appearing constitute novel environments in which wildlife must learn to coexist with people, thereby serving as ideal case studies with which to infer future human-wildlife interactions in shared landscapes.<em> </em>As a widely reviled and behaviorally plastic apex predator, the spotted hyena (<em>Crocuta crocuta</em>) is a model species for understanding how large carnivores navigate these human-caused 'landscapes of fear' in a changing world. Using high-resolution GPS collar data, we applied resource selection functions and step selection functions to assess spotted hyena landscape navigation and fine-scale movement decisions in relation to social-ecological features in a rapidly developing region comprising two protected areas: Lake Nakuru National Park and Soysambu Conservancy, Kenya. We then used camera trap imagery and Barrier Behavior Analysis (BaBA) to further examine hyena interactions with barriers. Our results show that environmental factors, linear infrastructure, human-carnivore conflict hotspots, and human tolerance were all important predictors for landscape-scale resource selection by hyenas, while human experience elements were less important for fine-scale hyena movement decisions. Hyena selection for these characteristics also changed seasonally and across land management types. Camera traps documented an exceptionally high number of individual spotted hyenas (234) approaching the national park fence at 16 sites during the study period, and BaBA results suggested that hyenas perceive protected area boundaries' semi-permeable electric fences as risky but may cross them out of necessity. Our findings highlight that the ability of carnivores to flexibly respond within human-caused landscapes of fear may be expressed differently depending on context, scale, and climatic factors. These results also point to the need to incorporate societal factors into multiscale analyses of wildlife movement to effectively plan for human-wildlife coexistence.</p>
Data from: Patterns of constitutive and induced herbivore defense are complex, but share a common genetic basis in annual and perennial monkeyflower
<p>Despite multiple ecological and evolutionary hypotheses that predict patterns of phenotypic relationships between plant growth, reproduction, and constitutive and/or induced resistance to herbivores, these hypotheses do not make any predictions about the underlying molecular genetic mechanisms that mediate these relationships. We investigated how divergent plant life-history strategies in the yellow monkeyflower and a life-history altering locus, DIV1 influence plasticity of phytochemical herbivory resistance traits in response to attack by two herbivore species with different diet breadth. Life-history strategy (annual vs. perennial) and the DIV1 locus significantly influenced levels of constitutive herbivory resistance, as well as resistance induction following both generalist and specialist herbivory. Perennial plants had higher total levels of univariate constitutive and induced defense than annuals, regardless of herbivore type. Annuals induced less in response to generalist herbivory than did perennials, while induction response was equivalent across the ecotypes for specialist herbivory. The effects of the DIV1 locus on levels of constitutive and induced defense were dependent on genetic background, the annual versus perennial haplotype of DIV1, and herbivore identity. The patterns of univariate induction due to DIV1 were non-additive and did not always match expectations based on patterns of divergence for annual/perennial parents. For example, perennial plants had higher levels of constitutive and induced defense than did annuals, but when the annual DIV1 was present in the perennial genetic background induction response to herbivory was higher than for the perennial parent lines. Patterns for multivariate defense arsenals generally echoed those of univariate, with annual and perennial monkeyflowers and those with alternative versions of DIV1 differing significantly in constitutive and induced resistance. Like univariate resistance, induced multivariate defense arsenals were affected by herbivore identity. Our results highlight the complexity of the genetic mechanisms underlying plastic response to herbivory. While a genetic locus underlying substantial phenotypic variation in life-history strategy and constitutive defense also influences defense plasticity, the induction response also depends on genetic background. This result demonstrates the potential for some degree of evolutionary independence between constitutive and induced defense, or induced defense and life-history strategy, in monkeyflowers.</p>
Walker et al - Prolactin and the shared regulation of parental care and cooperative helping behaviour in white-browed sparrow weaver societies - Data Set
<p>This file provides the data set supporting the analyses presented in a Walker et al manuscript titled "Prolactin and the shared regulation of parental care and cooperative helping behaviour in white-browed sparrow weaver societies" at first submission for review.</p>
FAIRmat Tutorial 4: NOMAD Oasis and FAIR data collaboration and sharing
<p>FAIRmat further develops NOMAD from a central publishing service to a federated data management platform. The NOMAD Oasis is part of this. Institutes, universities, and research groups use NOMAD Oasis as a local repository to manage their research data. Each method is different and requires ways of data acquisition, different data formats, different analysis tools, but FAIR-ness requires that all data is well described with rich specific metadata.</p> <p>In this tutorial, we focus on how to get started with NOMAD Oasis and adapt it to your research. One the first day, two talks will introduce you the general FAIRmat strategy and its "bottom-up" approach to manage heterogenous but FAIR data. On the second day, we will give the practical, step-by-step guides to get started with an Oasis: How you can install NOMAD Oasis, create example data, add schemas, and create ELNs.</p> <p> </p> <p><strong>Disclaimer:</strong> NOMAD is being continuously developed based on input and feedback from the scientific community. Hence the features, services or interface may have changed since the time of recording of this video. For up-to-date information please consult our latest tutorials and the NOMAD documentation <a href="https://nomad-lab.eu/prod/v1/docs/">https://nomad-lab.eu/prod/v1/docs/</a></p> <p><strong> </strong></p>
"Data Protection Can Sometimes Be a Nuisance" A Notification Study on Data Sharing Practices in City Apps
<p># A Notification Study on Data Sharing Practices in City Apps - Artifacts</p> <p>This archive contains the following artifacts of our study:<br>- [Pseudonymized data of our app analysis](pseudonymized_data.csv)<br> - a CSV file containing the data of our dynamic app analysis<br> - the columns represent the dates of our measurements<br> - the rows are the analyzed apps, we replaced the names<br> - the groups are as follows: <br> - 1 generic notification<br> - 2 legal notification<br> - 3 technical guidance notification<br> - each cell contains the trackers to which we observed HTTP requests, divided by `|`<br>- [Mail templates](./mail_templates/)<br> - the templates to the german notification mails we sent<br> - the `guides` directory contains the technical guidance specific to the observed tracker that we provided to the technical group</p>
The sharing of research raw data in journals indexed in the Cell & Tissue Engineering JCR category (2011-2015)
<p>The availability of research data sets is an important milestone since it can enhance the dynamics of research. This study aims to analyze the PubMed Central repository to determine the availability and type of raw data sets in Cell & Tissue Engineering journals indexed in the Journal Citation Reports. The number and types of files were registered. A search of the 21 journals from the Cell & Tissue Engineering category of the 2015 Journal Citation Reports was conducted. Information was collected from October to December 2016. A study of the supplementary material of the original articles published between 2011-2015 was performed through a search in the PubMed Central repository, which is the most used free full-text repository in biomedicine. Only articles with supplementary material were retrieved. The number and types of files were registered. In cases where a compressed file, such as a .zip or .rar file, was found, it was opened to check what kinds of files it contained.</p>
Data sharing in the life sciences : a study of researchers at the Norwegian University of Life Sciences
<p>Survey data collected in 2012 as basis for master thesis in library and information sciences titled "Data sharing in the life sciences : a study of researchers at the Norwegian University of Life Sciences" </p>
Research data sharing in Spain: exploring determinants, practices and perceptions
<p>Responses of the Spanish survey for research data, which was carried out within the framework of the project Datasea at the beginning of 2015. The purpose was to identify the habits and current experiences of Spanish researchers in health sciences in relation to the management and sharing of raw research data. The electronic questionnaire composed of 40 questions divided into three blocks that contained questions on the following aspects: A) Personal information; B) Creation and reuse of data; C) Preservation of data.</p>
CRAFT 2019 Shared Task data
<p>This data set consists of data used for the CRAFT Shared Task 2019.</p> <p>Version 3.1.3 of the CRAFT corpus was provided to participants as training data (CRAFT-3.1.3.tar.gz).</p> <p>During the evaluation phase, participants were provided the 30 plain text documents of the CRAFT evaluation set, along with ontologies used for concept annotation, other concept metadata files, and tokens required for the coreference resolution evaluation. (craft-st-2019-2019_test_data.tar.gz)</p> <p>Finally, the evaluation was completed using the gold standard annotation files available in evaluation-data.tar.gz.</p>
Data underlying the manuscript: "Analysis of Research Data Sharing in Scientific Articles on Climate Change in the Covid-19 Year. The Spanish case 2020".
<p>This is the research data for the manuscript "Analysis of Research Data Sharing in Scientific Articles on Climate Change in the Covid-19 Year. The Spanish case 2020".<br>The following is the original abstract: Introduction: Sharing research data on climate change would facilitate the development of solutions to curb its impact, for this, data needs to be shared in an optimal way. General objective: To identify how many Spanish scientific articles on climate change published during 2020 share their research data in some way. Specific objectives: a) Identify the attributes of shared research data b) Describe the characteristics of the case studies found on how research data are shared. Methodology: Qualitative and descriptive study analyzing nine attributes: availability (1), accessibility (2), format (3), license (4), linkage (5), funding (6), editorial policy (7), content (8), statistics (9). Results: We analyzed 2212 articles were analyzed, 1867 (84%) articles had no associated research data. The remaining 16% have associated research data: 152 (7%) articles deposited their data in repositories, 42 (2%) submitted their data as supplementary material, 136 (6%) will share their data upon request to the author and 15 (1%) do not have publication permissions. Conclusions: Researchers are willing to share their research data, but under different conditions. Researchers who reused research data did not share the new data they generated. There is a lack of training among researchers on how to manage their research data. There is information on the web on this topic, but it is not just a matter of publishing manuals, but also of creating training spaces within universities, institutes and research centers to build a community of researchers committed to Open Science.</p>
Data for "Information sharing within a social network is key to behavioral flexibility – lessons from mice tested under semi-naturalistic conditions"
<p>Data for "Information sharing within a social network is key to behavioral flexibility – lessons from mice tested under semi-naturalistic conditions", currently under review in Science Advances. </p>
Data and Code from: Dysregulation of zebrin-II cell subtypes is a shared feature across polyglutamine ataxia mouse models and human patients
<div> <div> <div> <p>Abstract</p> <p>Spinocerebellar ataxia type 7 (SCA7) is a genetic neurodegenerative disorder caused by a CAG- polyglutamine repeat expansion. Purkinje cells (PCs) are central to the pathology of ataxias, but their low abundance in the cerebellum underrepresents their transcriptomes in sequencing assays. To address this issue, we developed a PC enrichment protocol and sequenced individual nuclei from mice and patients with SCA7. Single-nucleus RNA sequencing in SCA7-266Q mice revealed dysregulation of cell identity genes affecting glia and PCs. Specifically, genes marking zebrin-II PC subtypes accounted for the highest proportion of DEGs in symptomatic SCA7-266Q mice. These transcriptomic changes in SCA7-266Q mice were associated with increased numbers of inhibitory synapses as quantified by immunohistochemistry and reduced spiking of PCs in acute brain slices. Dysregulation of zebrin-II cell subtypes was the predominant signal in PCs of SCA7-266Q mice and was associated with the loss of zebrin-II striping in the cerebellum at motor symptom onset. We furthermore demonstrated zebrin-II stripe degradation in additional mouse models of polyglutamine ataxia and observed decreased zebrin-II expression in cerebellum of patients with SCA7. Our results suggest that a breakdown of zebrin subtype regulation is a shared pathological feature of polyglutamine ataxias.</p> <p>Data and Code Availability</p> <p>Here you will find data and code associated with our manuscript "Dysregulation of zebrin-II cell subtypes is a shared feature across polyglutamine ataxia mouse models and human patients", Bartelt et al., <em>Sci. Trans. Med. </em>16, eadn5449 (2024).</p> <p>The data file labeled "HuCb_filtered.rds" is a processed and annotated single-nucleus RNA-seq Seurat object, containing the gene-level count data for the multiplexed snRNA-seq experiment performed on post-mortem human cerebellar tissues from patients with SCA7 and unaffected controls. Data obtained from WT and SCA7-266Q mice as described in our paper can be accessed in the NIH Gene Expression Omnibus under accession number GSE269430.</p> <p>There are three code files numbered 00 through 02 which contain analysis code for snRNA-seq data applied to both the mouse and human datasets. These files are sequential and will take the user from CellRanger output, to filtered and annotated Seurat objects, and include details for subclustering analysis as well as our pseudobulk DEseq2 differential expression approach. There are places where the user may need to modify the code based on their computer system, version of R or Seurat, and whether they are processing the 5 week, 8 week, or human data sets; these locations in the code are marked with comments.</p> <ul> <li>The first file, 00_Preprocessing_MULTIseq, begins with CellRanger filtered_feature_barcode_matrix output, extracts cell barcodes, utilizes the MULTIseq deMULTIplex software to match cell barcodes to oligo barcodes from MULTIseq fastq files, and annotates the Seurat file with metadata. Cell type identification and annotation also takes place in this file. Note: the deMULTIplex step will likely need to be run on a high performance compute cluster.</li> <li>The second file, 01_Seurat_Analysis, uses the filtered and annotated Seurat file to calculate useful QC metrics, investigate disease signals, and perform cell type subclustering analyses.</li> <li>The third file, 02_Pseudobulk_DEseq2, contains custom analysis code to extract raw counts for each cell type and each animal from the Seurat file, and uses the DEseq2 package to calculate DEGs, taking into account biological replicates, and raw read count differences between control and SCA7 animals.</li> </ul> </div> </div> </div>
Data from: Shifts to earlier selfing in sympatry may reduce costs of pollinator sharing
Coexisting plant congeners often experience strong competition for resources. Competition for pollinators can result in direct fitness costs via reduced seed set or indirect costs via heterospecific pollen transfer (HPT), causing subsequent gamete loss and unfit hybrid offspring production. Autonomous selfing may alleviate these costs, but to preempt HPT, selfing should occur early, before opportunities for HPT occur (i.e. "preemptive selfing hypothesis"). We evaluated conditions for this hypothesis in Collinsia sister species, C. linearis and C. rattanii. In field studies, we found virtually identical flowering times and pollinator sharing between congeners in sympatric populations. Compared to allopatric populations, sympatric C. linearis populations enjoyed higher pollinator visitation rates, whereas visitation to C. rattanii did not differ in sympatry. Importantly, the risk of HPT to each species in sympatry was strongly asymmetrical; interspecies visits comprised 40% of all flower-to-flower visits involving C. rattanii compared to just 4% involving C. linearis. Additionally, our greenhouse experiment demonstrated a strong cost of hybridization, when C. rattanii was the pollen donor. Together, these results suggest that C. rattanii pays the greatest cost of pollinator sharing. Matching predictions of the preemptive selfing hypothesis, C. rattanii exhibit significantly earlier selfing in sympatric relative to allopatric populations.
pGAN Synthetic Dataset: A Deep Learning Approach to Private Data Sharing of Medical Images Using Conditional GANs
<p>Synthetic dataset for <strong>A Deep Learning Approach to Private Data Sharing of Medical Images Using Conditional GANs</strong></p> <p><strong> Dataset specification:</strong></p> <ul> <li>MRI images of Vertebral Units labelled based on region</li> <li>Dataset is comprised of 10000 pairs of images and labels</li> <li>Image and label pair number k can be selected by: synthetic_dataset['images'][k] and synthetic_dataset['regions'][k]</li> <li>Images are 3D of size (9, 64, 64)</li> <li>Regions are stored as an integer. Mapping is 0: cervical, 1: thoracic, 2: lumbar</li> </ul> <p>Arxiv paper: <a href="https://arxiv.org/abs/2106.13199">https://arxiv.org/abs/2106.13199</a><br> Github code: <a href="https://github.com/tcoroller/pGAN/">https://github.com/tcoroller/pGAN/</a></p> <p>Abstract:</p> <p>Sharing data from clinical studies can facilitate innovative data-driven research and ultimately lead to better public health. However, sharing biomedical data can put sensitive personal information at risk. This is usually solved by anonymization, which is a slow and expensive process. An alternative to anonymization is sharing a synthetic dataset that bears a behaviour similar to the real data but preserves privacy. As part of the collaboration between Novartis and the Oxford Big Data Institute, we generate a synthetic dataset based on COSENTYX Ankylosing Spondylitis (AS) clinical study. We apply an Auxiliary Classifier GAN (ac-GAN) to generate synthetic magnetic resonance images (MRIs) of vertebral units (VUs). The images are conditioned on the VU location (cervical, thoracic and lumbar). In this paper, we present a method for generating a synthetic dataset and conduct an in-depth analysis on its properties of along three key metrics: image fidelity, sample diversity and dataset privacy.</p>
Shared motivations, goals and values in the practice of personal science - Qualitative data set
<p>269 transcribed excerpts coded from 22 interviews to self-researchers for the study "Shared motivations, goals and values in the practice of personal science - A community perspective on self-tracking for empirical knowledge". Interviews with participants were conducted via video conferencing and were based on a list of open-ended questions, separated into key sections around participation and collaboration in personal science. Participants who agreed to be interviewed, gave informed consent in like with the ethics approval by the Inserm Institutional Review Board (IRB) for this study, and regarding this data set, previous agreement in compliance with privacy and anonymity requirements. Academic article based on this dataset: Senabre Hidalgo, E., Ball, M. P., Opoix, M., & Greshake Tzovaras, B. (2022). Shared motivations, goals and values in the practice of personal science: a community perspective on self-tracking for empirical knowledge. <em>Humanities and Social Sciences Communications</em>, <em>9</em>(1), 1-12. <a href="https://doi.org/10.1057/s41599-022-01199-0">https://doi.org/10.1057/s41599-022-01199-0</a></p>
Cryptocurrency Fraud and Code Sharing Data Set and Analysis Code
<p>This release covers the state of the data and associated analysis code for determining code sharing between cryptocurrency codebases funded through the end of the original NSF CRII award. This material is based on work supported by the National Science Foundation under Grant CNS-1849729.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.