Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
247
datasets available to search
ShareScore release 0.7.1
Dataset results
247 results for “relatedness”
Data from: Using genetic relatedness to understand heterogeneous distributions of urban rat-associated pathogens
<p>Urban Norway rats (<i>Rattus norvegicus</i>) carry several pathogens transmissible to people. However, pathogen prevalence can vary across fine spatial scales (i.e., by city block). Using a population genomics approach, we sought to describe rat movement patterns across an urban landscape, and to evaluate whether these patterns align with pathogen distributions. We genotyped 605 rats from a single neighborhood in Vancouver, Canada and used 1,495 genome-wide single nucleotide polymorphisms to identify parent-offspring and sibling relationships using pedigree analysis. We resolved 1,246 pairs of relatives, of which only 1% of pairs were captured in different city blocks. Relatives were primarily caught within 33 meters of each other leading to a highly leptokurtic distribution of dispersal distances. Using binomial generalized linear mixed models we evaluated whether family relationships influenced rat pathogen status with the bacterial pathogens <i>Leptospira interrogans</i>, <i>Bartonella tribocorum</i>, and <i>Clostridium difficile</i>, and found that an individual's pathogen status was not predicted any better by including disease status of related rats. The spatial clustering of related rats and their pathogens lends support to the hypothesis that spatially restricted movement promotes the heterogeneous patterns of pathogen prevalence evidenced in this population. <span>Our findings also highlight the utility of evolutionary tools to understand movement and rat-associated health risks in urban landscapes.</span></p>
Human and Machine Judgements for Russian Semantic Relatedness
<p>Semantic relatedness of terms represents similarity of meaning by a numerical score. On the one hand, humans easily make judgements about semantic relatedness. On the other hand, this kind of information is useful in language processing systems. While semantic relatedness has been extensively studied for English using numerous language resources, such as associative norms, human judgements and datasets generated from lexical databases, no evaluation resources of this kind have been available for Russian to date. Our contribution addresses this problem. We present five language resources of different scale and purpose for Russian semantic relatedness, each being a list of triples (wordi, wordj , similarityij ). Four of them are designed for evaluation of systems for computing semantic relatedness, complementing each other in terms of the semantic relation type they represent. These benchmarks were used to organise a shared task on Russian semantic relatedness, which attracted 19 teams. We use one of the best approaches identified in this competition to generate the fifth high-coverage resource, the first open distributional thesaurus of Russian. Multiple evaluations of this thesaurus, including a large-scale crowdsourcing study involving native speakers, indicate its high accuracy.</p> <p>For more details see: </p> <ul> <li>The web page of the RUSSE evaluation campaign: http://russe.nlpub.ru/downloads</li> <li>The original publication "Panchenko A., Ustalov D., Arefyev N., Paperno D. Konstantinova N., Loukachevitch N. and Biemann C. undefinedHuman and Machine Judgements about Russian Semantic Relatedness. In Proceedings of the 5th Conference on Analysis of Images, Social Networks and Texts (AIST'2016). Communications in Computer and Information Science (CCIS). Springler-Verlag Berlin Heidelberg": https://www.lt.informatik.tu-darmstadt.de/fileadmin/user_upload/Group_LangTech/publications/aist_2016_hmj.pdf</li> </ul>
Effectively controlling for sample relatedness in large-scale GWAS: application to 79 EHR-derived longitudinal traits
<p><span>Sample relatedness is a major confounder in genome-wide association studies (GWAS), potentially leading to inflated type I error rates if not appropriately controlled. A common strategy is to incorporate a random effect related to genetic relatedness matrix (GRM) into regression models. However, this approach is challenging for large-scale GWAS of complex traits, such as longitudinal traits. Here we propose a scalable and accurate analysis framework, SPA<sub>GRM</sub>, which controls for sample relatedness via a precise approximation of the joint distribution of genotypes. SPA<sub>GRM</sub> can utilize GRM-free models and thus is applicable to various trait types and statistical methods, including linear mixed models and generalized estimation equations for longitudinal traits. A hybrid strategy incorporating saddlepoint approximation greatly increases the accuracy to analyze low-frequency and rare genetic variants, especially in unbalanced phenotypic distributions. We also introduce SPA<sub>GRM(CCT)</sub> to aggregate the results following different models via Cauchy combination test. Extensive simulations and real data analyses demonstrated that SPA<sub>GRM</sub> maintains well-controlled type I error rates and SPA<sub>GRM(CCT)</sub> can serve as a broadly effective method. Applying SPA<sub>GRM</sub> to 79 longitudinal traits extracted from</span><span> </span><span>UK Biobank primary care data, we identified 7,463 genetic loci, making a pioneering attempt to conduct GWAS for these traits as longitudinal traits.</span></p>
The impact of species phylogenetic relatedness on invasion varies distinctly along resource versus nonresource environmental gradients
<p><span>Understanding why certain plant communities are vulnerable to alien invasive species is essential to predicting and controlling invasion in a changing environment. Darwin's naturalization hypothesis suggests that non-native species should be more successful in communities where their close relatives are absent. Empirical tests of this hypothesis, however, have produced mixed results. Using plot-level data from natural forests along elevational transects covering strong environmental gradients, we examined whether the invasion of the globally invasive species <em>Ageratina adenophora</em> can be explained by environmental filtering and/or competition from closely related species linked to environmental gradients. Abundant precipitation, warm temperatures, open canopies, and postfire environments facilitated <em>A. adenophora</em> invasion, whereas resident taxonomic richness suppressed its invasion. Importantly, we found that invader-resident relatedness had a strong negative effect on invader cover under resource scarcity conditions (e.g., low water availability), but not under nonresource environmental stress conditions (e.g., low temperature). Our findings help reconcile the varied applicability of Darwin's naturalization hypothesis to biological invasions in a changing world.</span></p>
The SICK (Sentences Involving Compositional Knowledge) dataset for relatedness and entailment
<p>The SICK data set consists of about 10,000 English sentence pairs, generated starting from two existing sets: the <a href="http://nlp.cs.illinois.edu/HockenmaierGroup/data.html">8K ImageFlickr data set</a> and the <a href="http://www.cs.york.ac.uk/semeval-2012/task6/index.php?id=data">SemEval 2012 STS MSR-Video Description data set</a>. We randomly selected a subset of sentence pairs from each of these sources and we applied a 3-step generation process: first, the original sentences were normalized to remove unwanted linguistic phenomena; the normalized sentences were then expanded to obtain up to three new sentences with specific characteristics suitable to CDSM evaluation; as a last step, all the sentences generated in the expansion phase were paired with the normalized sentences in order to obtain the final data set.</p> <p>Each sentence pair was annotated for relatedness and entailment by means of crowdsourcing techniques. The <strong>sentence relatedness score</strong> (on a 5-point rating scale) provides a direct way to evaluate CDSMs, insofar as their outputs are meant to quantify the degree of semantic relatedness between sentences; the categorizations in terms of the <strong>entailment relation between the two sentences</strong> (with <em>entailment, contradiction</em>, and <em>neutral</em> as gold labels) is also a crucial aspect to consider, since detecting the presence of entailment is one of the traditional benchmarks of a successful semantic system.</p> <p>In the final set, gold scores for relatedness and entailment were distributed as follows: the relatednes scoring resulted in 923 pairs within the [1,2) range, 1373 pairs within the [2,3) range, 3872 pairs within the [3,4) range, and 3672 pairs within the [4,5] range; the entailment annotation led to 5595 <em>neutral</em> pairs, 1424 <em>contradiction</em> pairs, and 2821 <em>entailment</em> pairs.</p> <p><strong>Files</strong></p> <ul> <li>SICK.zip (main file)</li> <li>SICK_Annotated.zip (a version of the data set annotated for the expansion rule which was used in each case)</li> <li>SICK_subsets.zip (a Indexes specifying further classifications, used in the JLRE 2016 publication)</li> </ul> <p> </p>
Data & Analysis Script for: Phylogenetic relatedness to native congeners drives insect abundance and diversity hosted by non-native trees
<p>The dataset contains all necessary data to reproduce the findings presented in Schweiger et al. 2023 - Phylogenetic relatedness to native congeners drives insect abundance and diversity hosted by non-native trees (submitted).</p> <p>The code necessary to reproduce the findings is included within this repository. The code contains comments. Please note, if you want to reproduce the findings you will have to change file path information matching your personal computer to be able to re-run the code.</p> <p>This data includes the biodiversity raw data collected for the manuscript. It <strong>does not </strong>include data used to calculate geographic, climatic or phylogenetic distances, as these data are freely available and necessary information to reproduce calculations are given within the Material & Methods section.</p> <p>All data is provided within one Excel file. Please, pay attention to the provided ReadMe sheet containing metadata information on the dataset.</p> <p>Please carefully read provided information within ReadMe, Metadata and Code description.</p>
Figure 3 in Effects of genetic relatedness, spatial distance, and context on intraspecific aggression in the red wood ant Formica pratensis (Hymenoptera: Formicidae)
Figure 3. Correlation between spatial distance and aggression levels in the field. Open circles correspond to monodomous colonies and filled circles correspond to the polydomous one.
Figure 1. Map showing the localities where F in Effects of genetic relatedness, spatial distance, and context on intraspecific aggression in the red wood ant Formica pratensis (Hymenoptera: Formicidae)
Figure 1. Map showing the localities where F. pratensis colonies were sampled for the analysis of genetic relatedness and tested for their aggressive behavior towards each other. The numbers denote the localities. 1: Balaban village (N 41°49ʹ18ʺ, E 27°40ʹ44ʺ) containing three nests; B1, B2, and B3, 2: Asilbeyli village (N 41°39ʹ32ʺ, E 27°13ʹ50ʺ), one nest (As), 3: Ulukonak village (N 41°39ʹ35ʺ, E 27°01ʹ52ʺ) one nest (U), 4: Doğanköy village (N 41°56ʹ12ʺ, E 26°41ʹ20ʺ) one nest (D), and 5: Ahmetler village (N 42°00ʹ37ʺ, E 27°11ʹ12ʺ), three nests; Ah1, Ah2, and Ah3.
Data from: Phylogenetic relatedness drives protists assembly in marine and terrestrial environments
<p>Aim: Assembly of protists communities is known to be driven mainly by environmental filtering, but the imprint of phylogenetic relatedness is unknown. In this study, we aim to test the degree at which co-occurrences and co-exclusions of protists in different phylogenetic relatedness classes are deviating from random expectation in two ecosystems in order to link them to ecological processes.</p> <p>Location: Global open-oceans and Neotropical rainforest soils</p> <p>Major taxa: Protists</p> <p>Time period: 2009-2013</p> <p>Methods: Protist metabarcoding data originated from two large scale studies. Co-occurrence and co-exclusion networks were constructed using a recent method combining a null distribution model with Spearman's rank correlation coefficients among pairs of OTU. Phylogenetic relatedness was estimated using either global pairwise sequence distance or phylogenetic distance inferred from best maximum-likelihood trees derived from multiple alignments of OTU representative sequences. Significance of observed patterns relating networks and phylogenies were evaluated by distance classes against two null models in which either the tips of the phylogenetic trees or the network edges were randomized.</p> <p>Results: Closely-related protists co-occurred more often than expected by chance in all datasets, but also co-excluded less often than expected by chance in the marine dataset only. Concurrent excess of co-occurrences and co-exclusions were observed at intermediate phylogenetic distances in the marine dataset.</p> <p>Main conclusions: This suggest that environmental filtering and dispersal limitation are the dominant forces driving protists co-occurrences in both environments, while signal of competitive exclusion was only detected in the marine environment. Co-exclusion differences are potentially linked to the individual environments: marine waters are more homogeneous, while the rainforest soils contain a myriad of nutrient rich micro-environment reducing the strength of mutual exclusion.</p>
Cloacal microbiomes of sympatric and allopatric Sceloporus lizards vary with environment and host relatedness
<p><span>Animals and their microbiomes exert reciprocal influence; the host's environment, physiology, and phylogeny can impact the composition of the microbiome, while the microbes present can affect host behavior, health, and fitness. While some microbiomes are highly malleable, specialized microbiomes that provide important functions can be more robust to environmental perturbations. Recent evidence suggests <em>Sceloporus</em> <em>virgatus</em> has one such specialized microbiome, which functions to protect eggs from fungal pathogens during incubation. Here, we examine the cloacal microbiome of three different <em>Sceloporus</em> species (spiny lizards; Family Phrynosomatidae) – <em>Sceloporus</em> <em>virgatus</em>, <em>Sceloporus</em> <em>jarrovii</em>, and <em>Sceloporus</em> <em>occidentalis</em>. We compare two species with different reproductive modes (oviparous vs. viviparous) living in sympatry: <em>S</em>. <em>virgatus</em> and <em>S</em>. <em>jarrovii</em>. We compare sister species living in similar habitats (riparian oak-pine woodlands) but different latitudes: <em>S</em>. <em>virgatus</em> and <em>S</em>. <em>occidentalis</em>. And, we compare three populations of one species (<em>S</em>. <em>occidentalis</em>) living in different habitat types: beach, low-elevation forest, and the riparian woodland. We found differences in beta diversity metrics between all three comparisons, although those differences were more extreme between animals in different environments, even though those populations were more closely related. Similarly, alpha diversity varied among the <em>S</em>. <em>occidentalis</em> populations and between <em>S</em>. <em>occidentalis</em> and <em>S</em>. <em>virgatus</em>, but not between sympatric <em>S</em>. <em>virgatus</em> and <em>S</em>. <em>jarrovii</em>. Despite these differences, all three species and all three populations of <em>S</em>. <em>occcidentalis</em> had the same dominant taxon, <em>Enterobacteriaceae</em>. The majority of the variation between groups was in low abundance taxa and at the ASV level, and responded to habitat differences, geographic distance, and host relatedness. Understanding wild microbiomes and factors that influence their composition is important to understanding the ecology and evolution of the host animals. </span></p>
Data from: Genomic data reveal unexpected relatedness between a brown female Eastern bluebird and her brood
<p>Because plumage coloration is frequently involved in sexual selection, for both male and female mate choice, birds with aberrant plumage should have fewer mating opportunities and thus lower reproductive output. Here we report an Eastern Bluebird (<em>Sialia sialis</em>) female with a brown phenotype that raised a brood of four chicks to fledging. The brown female and her mate were only related to their social offspring to the second degree and one of the offspring was a half-sibling. We propose four family tree scenarios and discuss their implications (e.g., extra-pair paternity, conspecific brood parasitism), but regardless of the tree, the brown female was able to find a mate, which may have been facilitated by the bottleneck created by the severe snowstorms in February 2021.</p>
The impact of species phylogenetic relatedness on invasion varies distinctly along resource versus nonresource environmental gradients
Open the record for dataset details and reuse information.
Cloacal microbiomes of sympatric and allopatric Sceloporus lizards vary with environment and host relatedness
Open the record for dataset details and reuse information.
Data from: Altruism or selfishness: Floral behavior based on genetic relatedness with neighboring plants
Open the record for dataset details and reuse information.
Data from: Inbreeding and competitor’s genetic relatedness affect dynamic male color-ornament expression in a cichlid fish
Open the record for dataset details and reuse information.
Data from: Genomic data reveal unexpected relatedness between a brown female Eastern bluebird and her brood
Open the record for dataset details and reuse information.
Data from: Phylogenetic relatedness drives protists assembly in marine and terrestrial environments
Open the record for dataset details and reuse information.
Going with the flow? Relative importance of riverine hydrologic connectivity versus tidal influence for spatial structure of genetic diversity and relatedness in a foundational submersed aquatic plant
Open the record for dataset details and reuse information.
Data from: Using genetic relatedness to understand heterogeneous distributions of urban rat-associated pathogens
Open the record for dataset details and reuse information.
Genome–scale approach to study the genetic relatedness among Brucella melitensis strains - wgMLST schema for Brucella melitensis
<p><strong>wgMLST schema for <em>Brucella melitensis</em></strong></p> <p> </p> <p><strong>Schema creation</strong></p> <p>The wgMLST schema was created using the 60 complete genomes of <em>Brucella melitensis </em>available at <a href="https://enterobase.warwick.ac.uk/species/index/ecoli">NCBI</a>, as of January 2019, with the chewBBACA v2.0.11 suite (<a href="https://github.com/B-UMMI/chewBBACA">https://github.com/B-UMMI/chewBBACA</a>), using a training file generated by Prodigal v2.6.3 from the <em>B. melitensis</em> 16M reference genome (RefSeq Accession NC_003317 and NC_003318). For curation and validation, the wgMLST schema was further populated with 212 additional draft genomes: 157 draft genomes (downloaded from NCBI in January 2019) and 55 draft genomes assembled with<a href="https://github.com/B-UMMI/INNUca"> INNUca v3.1</a> (PRJEB30030).</p> <p>File 'Bmelitensis_wgMLST_2656_schema.tar.gz' contains the wgMLST schema formatted for chewBBACA and includes a total of 2656 loci.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.