Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

227

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

227 results for “demographic history”

Learn how ShareScore rates datasets ↗
zenodo44/100

Unique demographic history and population substructure among the Coorgs of Southern India

<p>Quality filtered GSA data of the individuals analysed in Mukhopadhyay et al., 2024 from Coorg, Karnataka, India.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

1805-1898 Census Records of Lausanne : a Long Digital Dataset for Demographic History

<p><strong>Context. </strong>This historical dataset stems from the project of automatic extraction of 72 census records of Lausanne, Switzerland. The complete dataset covers a century of historical demography in Lausanne (1805-1898), which corresponds to 18,831 pages, and nearly 6 million cells.</p> <p><strong>Content.</strong> The data published in this repository correspond to a first release, i.e. a diachronic slice of one register every 8 to 9 years. Unfortunately, the remaining data are currently under embargo. Their publication will take place as soon as possible, and at the latest by the end of 2023. In the meantime, the data presented here correspond to a large subset of 2,844 pages, which already allows to investigate most research hypotheses.</p> <p><strong>Description. </strong>The population censuses, digitized by the <a href="https://www.lausanne.ch/vie-pratique/culture/bibliotheques-et-archives/archives.html">Archives of the city of Lausanne</a>, continuously cover the evolution of the population in Lausanne throughout the 19th century, starting in 1805, with only one long interruption from 1814 to 1831. Highly detailed, they are an invaluable source for studying migration, economic and social history, and traces of cultural exchanges not only with Bern, but also with France and Italy. Indeed, the system of tracing family origin, specific to Switzerland, allows to follow the migratory movements of families long before the censuses appeared. The bourgeoisie is also an essential economic tracer. In addition, censuses extensively describe the organization of the social fabric into family nuclei, around which gravitate various boarders, workers, servants or apprentices, often living in the same apartment with the family.</p> <p><strong>Production. </strong>The structure and richness of censuses have also provided an opportunity to develop automatic methods for processing structured documents. The processing of censuses includes several steps, from the identification of text segments to the restructuring of information as digital tabular data, through Handwritten Text Recognition and the automatic segmentation of the structure using neural networks. Please note that the detailed extraction methodology, as well as the complete evaluation of performance and reliability is published in:</p> <ul> <li>Petitpierre R., Rappo L., Kramer M. (2023). <em>An end-to-end pipeline for historical censuses processing</em>. International Journal on Document Analysis and Recognition (IJDAR). doi: <a href="https://doi.org/10.1007/s10032-023-00428-9">10.1007/s10032-023-00428-9</a></li> </ul> <p><strong>Data structure.</strong> The data are structured in rows and columns, with each row corresponding to a household. Multiple entries in the same column for a single household are separated by vertical bars &lang;|&rang;. The center point &lang;&middot;&rang; indicates an empty entry. For some columns (e.g., street name, house number, owner name), an empty entry indicates that the last non-empty value should be carried over. The page number is in the last column.</p> <p><strong>Liability. </strong>The data presented here are not curated nor verified. They are the raw results of the extraction, the reliability of which was thoroughly assessed in the above-mentioned publication. We insist on the fact that for any reuse of this data for research purposes, the implementation of an appropriate methodology is necessary. This may typically include string distance heuristics, or statistical methodologies to deal with noise and uncertainty.</p>

opencc-by-4.0Mar 2023View details →
dryad40/100

Data from: Evolutionary and demographic history of the Californian scrub white oak species complex: an integrative approach

<p>Understanding the factors promoting species formation is a major task in evolutionary research. Here, we employ an integrative approach to study the evolutionary history of the Californian scrub white oak species complex (genus <em>Quercus</em>). To infer the relative importance of geographical isolation and ecological divergence in driving the speciation process, we (i) analyzed inter- and intra-specific patterns of genetic differentiation and employed an approximate Bayesian computation (ABC) framework to evaluate different plausible scenarios of species divergence. In a second step, we (ii) linked the inferred divergence pathways with current and past species distribution models, and (iii) tested for niche differentiation and phylogenetic niche conservatism across taxa. ABC analyses showed that the most plausible scenario is the one considering the divergence of two main lineages followed by a more recent pulse of speciation. Genotypic data in conjunction with species distribution models and niche differentiation analyses support that different factors (geography vs. environment) and modes of speciation (parapatry, allopatry and maybe sympatry) have played a role in the divergence process within this complex. We found no significant relationship between genetic differentiation and niche overlap, which probably reflects niche lability and/or that multiple factors have contributed to speciation. Our study shows that different mechanisms can drive divergence even among closely related taxa representing early stages of species formation and exemplifies the importance of adopting integrative approaches to get a better understanding of the speciation process.</p>

opencc-zeroDec 2014View details →
zenodo40/100

Demographic history and natural selection shape patterns of deleterious mutation load and barriers to introgression across Populus genome

<p><br> Abbreviation of species names in each folder: Palb, P. alba; Pade, P. adenopoda; Pdav, P. davidiana; Ptra, P. tremula; Ptrs, P. tremuloides; Prot, P. rotundifolia; Pqio,P. qiongdaoensis.</p> <p>1. FST<br> Relative divergence (FST) for pairwise species comparisons was calculated for all sites with 100 Kbp non-overlapping windows.&nbsp;</p> <p>2. dxy<br> Absolute divergence (dxy) was calculated for all sites with 100 Kbp non-overlapping windows.&nbsp;</p> <p>3. Nucleotide diversity<br> Nucleotide diversity (&pi;) was calculated for all sites with 100 Kbp non-overlapping windows.&nbsp;</p> <p>4. Derived allele frequency<br> The derived frequencies of 4 different functional categories. Each folder contains seven Populus resluts</p> <p>5. Derived_allele_statistics<br> The statistics of homozygous and &nbsp;heterozygous derived alleles for loss of function, deleterious, tolerated and synonymous variants for each individual. The last two individuals in each file are outgroups&nbsp;</p> <p>6. dsuite-dinvestigate<br> The outputs of 10 trios using program Dinvestigate from Dsuite. The sliding window is 50 SNPs, and the step is 20 SNPs.</p> <p>7. Recombination rate<br> The result of population-scaled recombination rate was calculated by LDhat v2.2.</p> <p>8. Volcanofinder<br> Genome-wide scans of introgression sweeps within each species was implemented using VolcanFinder v.1.0 with the Model over 10 Kbp non-overlapping windows.</p> <p>9. ihh12<br> phased SNPs were used to computed ihh12 by selscan v1.3.0.&nbsp;</p> <p>10 populus162.phased.recode.vcf.gz<br> SNPs were phased with Beagle v.4.1 for the 162 non-hybrid individuals.</p> <p>11 populus227.snp.rm_indel.para_filter.biallelic.GQ30.max_miss20.bed.recode.vcf.gz&nbsp;<br> The vcf of 227 Populus samples.&nbsp;</p>

opencc-by-4.0Nov 2021View details →
dryad40/100

Phylogenomics, introgression, and demographic history of South American true toads (Rhinella)

<p>The effects of genetic introgression on species boundaries and how they affect species' integrity and persistence over evolutionary time have received increased attention. The increasing availability of genomic data has revealed contrasting patterns of gene flow across genomic regions, which impose challenges to inferences of evolutionary relationships and of patterns of genetic admixture across lineages. By characterizing patterns of variation across thousands of genomic loci in a widespread complex of true toads (<em>Rhinella</em>), we assess the true extent of genetic introgression across species thought to hybridize to extreme degrees based on natural history observations and multi-locus analyses. Comprehensive geographic sampling of five large-ranged Neotropical taxa revealed multiple distinct evolutionary lineages that span large geographic areas and, at times, distinct biomes. The inferred major clades and genetic clusters largely correspond to currently recognized taxa; however, we also found evidence of cryptic diversity within taxa. While previous phylogenetic studies revealed extensive mito-nuclear discordance, our genetic clustering analyses uncovered several admixed individuals within major genetic groups. Accordingly, historical demographic analyses supported that the evolutionary history of these toads involved cross-taxon gene flow both at ancient and recent times. Lastly, ABBA-BABA tests revealed widespread allele sharing across species boundaries, a pattern that can be confidently attributed to genetic introgression as opposed to incomplete lineage sorting. These results confirm previous assertions that the evolutionary history of <em>Rhinella</em> was characterized by various levels of hybridization even across environmentally heterogeneous regions, posing exciting questions about what factors prevent complete fusion of diverging yet highly interdependent evolutionary lineages.</p>

opencc-zeroNov 2021View details →
dryad40/100

Joint analysis of microsatellites and flanking sequences enlightens complex demographic history of interspecific gene flow and vicariance in rear-edge oak populations

<p><span>Inference of recent population divergence requires fast evolving markers and necessitates to differentiate shared genetic variation caused by ancestral polymorphism and gene flow. Theoretical research shows that the use of compound marker systems integrating linked polymorphisms with different mutational dynamics, such as a microsatellite and its flanking sequences, can improve estimation of population structure and inference of demographic history, especially in the case of complex population dynamics. However, empirical application in natural populations has so far been limited by lack of suitable methods for data collection. A solution comes from the development of sequence-based microsatellite genotyping which we used to study molecular variation at 36 sequenced nuclear microsatellites in seven <em>Quercus canariensis</em> and four <em>Q. faginea</em> rear-edge populations across Algeria. We aim to decipher their taxonomic relationship, past evolutionary history and recent demographic trajectory. First, we compare the estimation of population genetics parameters and simulation-based inference of demographic history from microsatellite sequence alone, flanking sequence alone or the combination of linked microsatellite and flanking sequence variation. Second, we apply random forest approximate Bayesian computation to identify which of these sequence types is most informative. Whereas analysing microsatellite variation alone indicates recent interspecific gene flow, additional information gained by integrating nucleotide variation in flanking sequences, by reducing homoplasy, suggests ancient interspecific gene flow followed by drift in isolation instead. The weight of each polymorphism in the inference also demonstrates the value of linked variations with contrasted mutation dynamic to improve estimation of both demographic and mutational parameters.</span></p>

opencc-zeroJun 2022View details →
zenodo40/100

Indirect genetic effects are shaped by demographic history and ecology in Arabidopsis thaliana

<p><em>This folder contains data &amp; code used for the study &quot;Indirect genetic effects are shaped by demographic history and ecology in Arabidopsis thaliana&quot;</em></p> <p>All data analyzed in the study are stored in the folder &quot;data&quot;:</p> <ul> <li>&quot;pheno_file.csv&quot;: the main phenotypic file corresponding to the experiment with paired plants used to estimate Indirect Genetic Effects.</li> <li>&quot;pheno_file_single_plants.csv&quot;: phenotypic file with measurements of plant biomasses in the absence of competition (single plants)</li> <li>&quot;call_method_75_TAIR9.csv&quot;: genomic data (SNPs) for each accession from the RegMap panel (ref [1])</li> <li>&quot;Data_geo_RegMap_accessions.csv&quot;: geographic localization of each accession from the RegMap panel (ref [2])</li> <li>&quot;igeGWAS_scores.csv&quot;: Genome-Wide Association Study (GWAS) results reporting for each SNP from the RegMap panel the p-value and estimated effect sizes of their direct and indirect genetic effects</li> <li>&quot;1001_accessions_info.csv&quot;: geographic localization and admixture group for each accession from the 1001 genomes project (ref [3])</li> <li>&quot;snp_data_all_samples.txt&quot;: allelic value of each accession from the 1001 genomes project at the eleven top SNPs associated with IGE</li> <li>&quot;sample_names.txt&quot;: names of&nbsp; accessions listed in the file &quot;snp_data_all_samples.txt&quot;</li> <li>&quot;climatic_data.csv&quot;: climatic data for each accessions from the 1001 genomes project (ref [4])</li> <li>&quot;candidate_genes_all.csv&quot;:&nbsp; list of all genes (and associated GO terms) with a non-synonymous, nonsense, or frameshift mutation in close proximity (distance &lt; half LD decay distance) and high linkage (r2&gt;0.5) with a SNP significantly associated with IGE</li> <li>&quot;genes.coord.bed&quot;: list of all genes in a +- 500 kb around top IGE SNPs and their coordinates</li> <li>&quot;AllGenes_fst.GeneID.txt&quot;: pairwise Fst computed between each pair of admixture groups, for all genes annotated in the genome of A. thaliana</li> </ul> <p>&quot;ABBA_BABA&quot; subfolder contains ABBA_BABA statistics computed for each individual chromosome (Chr1-Chr5) using genomic windows of 20 kb with at least 250 SNPs per windows. ABBA-BABA statistics were computed using custom python scripts from https://github.com/simonhmartin/genomics_general</p> <p><br> &quot;GEA&quot; subfolder contains Genome-Environment Association results, with one file per chromosome x climatic variable. Climatique variable are indexed, following the order listed in the file &quot;Climatic_variables.txt&quot; within the subfolder &quot;GEA&quot;. GEA analysis were run with the gemma program: https://github.com/genetics-statistics/GEMMA.</p> <p><br> &quot;LD_IGE_SNPs&quot; subfolders contains the list of SNPs located at +- 2Mb of a significant IGE SNP (one file per IGE SNP, named &quot;SNPalias_LDSimplified.csv&quot;) and their linkage (r2) with the IGE SNP. It also contains the file &quot;LD_windows_sizes.csv&quot; with the half LD decay distances for all significant IGE SNP.</p> <p>All analysis performed to produce the tables and figures presented in the study (main manuscript &amp; supplementary information) were done with the R script &quot;Arabidopsis_IGE_analysis.R&quot;, which uses &quot;manhattan_custom.R&quot; as a source function to produce custom manhattan plots.</p> <p>&nbsp;</p> <p><strong>REFERENCES:</strong></p> <p>[1] Horton MW, Hancock AM, Huang YS, Toomajian C, Atwell S, Auton A, Muliyati NW, Platt A, Sperone FG, Vilhj&aacute;lmsson BJ, et al. 2012. Genome-wide patterns of genetic variation in worldwide Arabidopsis thaliana accessions from the RegMap panel. Nature Genetics 44: 212&ndash;216.</p> <p>[2] Anastasio AE, Platt A, Horton M, Grotewold E, Scholl R, Borevitz JO, Nordborg M, Bergelson J. 2011. Source verification of mis-identified Arabidopsis thaliana accessions. The Plant Journal 67: 554&ndash;566.</p> <p>[3] 1001 Genomes Consortium. 2016. 1,135 genomes reveal the global pattern of polymorphism in Arabidopsis thaliana. Cell 166: 481&ndash;491.</p> <p>[4] Ferrero-Serrano &Aacute;, Assmann SM. 2019. Phenotypic and genome-wide association with the local environment of Arabidopsis. Nature Ecology &amp; Evolution 3: 274&ndash;285.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Figure 2. Bayesian posterior probability 50 in Population genetic structure and demographic history of the Chinese endemic Mongoloniscus sinensis (Dollfus, 1901) (Isopoda: Oniscidea)

Figure 2. Bayesian posterior probability 50% majority-rule consensus tree of the M. sinensis haplotypes. Out-group was Ligia occidentalis; the map showed mitochondrial haplotype clades of Porcellio gigliotose and Trachelipus semiproiectus. The numbers above joints are the bootstrap support values of MP value, ML value, and the posterior probabilities of the BI tree, respectively (MP/ML/BI).

opencc-by-4.0Dec 2016View details →
zenodo40/100

Figure 1 in Population genetic structure and demographic history of the Chinese endemic Mongoloniscus sinensis (Dollfus, 1901) (Isopoda: Oniscidea)

Figure 1. Locations of sampled populations and geographical distribution of Mongoloniscus sinensis mtDNA clades. Mountain ranges are in green letters as follows: H—Heng Mountain, TH—Taihang Mountain, WT—Wutai Mountain, LL—Lvliang Mountain, TY–Taiyue Mountain, ZT—Zhongtiao Mountain.

opencc-by-4.0Dec 2016View details →
zenodo40/100

Figure 3 in Population genetic structure and demographic history of the Chinese endemic Mongoloniscus sinensis (Dollfus, 1901) (Isopoda: Oniscidea)

Figure 3. Parsimonious network of M. sinensis haplotypes. Each circle represents a haplotype, with the area of the circle proportional to its frequency. The evolutionary clades C1 - C6 are shown in different colors, and median vector (mvl-mv3) is indicated in red.

opencc-by-4.0Dec 2016View details →
zenodo40/100

Figure 4 in Genetic diversity, population structure and demographic history of Dugesia japonica in Taihang Mountains

Figure 4. Median-joining haplotype network based on mitochondrial gene COI. The four ellipses represent four clades in Figure 3, respectively. Each circle represents a haplotype, the area of the circle is proportional to the frequency of haplotypes, and black dots represent hypothetical unobserved haplotypes. Different populations are shown in different colors.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 3 in Genetic diversity, population structure and demographic history of Dugesia japonica in Taihang Mountains

Figure 3. Maximum likelihood (ML) and Bayesian inference (BI) phylogentic trees based on mitochondrial gene COI. Dugesia ryukyuensis (Genbank accession no. AB618488) serves as the outgroup. The broken lines denote inconsistent branches. Bootstrap percentages (BP,&gt;50 only) of ML analysis and posterior probabilities (PP,&gt;0.50 only) of Bayesian inference are shown above and below the branch, respectively. HG—haplogroup.

opencc-by-4.0Dec 2021View details →
dryad40/100

Whole genome demographic models indicate divergent effective population size histories shape contemporary genetic diversity gradients in a montane bumble bee

<p>Understanding historical range shifts and population size variation provides important context for interpreting contemporary genetic diversity. Methods to predict changes in species distributions and model changes in effective population size (N<sub>e</sub>) using whole genomes make it feasible to examine how temporal dynamics influence diversity across populations. We investigate N<sub>e</sub> variation and climate-associated range shifts to examine the origins of a previously observed latitudinal heterozygosity gradient in the bumble bee <em>Bombus</em> <em>vancouverensis</em> Cresson (Hymenoptera: Apidae: <em>Bombus</em> Latreille) in western North America. We analyze whole genomes from a latitude-elevation cline using sequentially Markovian coalescent models of N<sub>e</sub> through time to test whether relatively low diversity in southern high-elevation populations is a result of long-term differences in N<sub>e</sub>. We use Maxent models of the species range over the last 130,000 years to evaluate range shifts and stability. N<sub>e</sub> fluctuates with climate across populations, but more genetically diverse northern populations have maintained greater Ne over the late Pleistocene and experienced larger expansions with climatically favorable time periods. Northern populations also experienced larger bottlenecks during the last glacial period which matched the loss of range area near these sites, however, bottlenecks were not sufficient to erode diversity maintained during periods of large N<sub>e</sub>. A genome sampled from an island population indicated a severe postglacial bottleneck, indicating that large recent post-glacial declines are detectable if they have occurred. Genetic diversity was not related to niche stability or glacial-period bottleneck size. Instead, spatial expansions and increased connectivity during favorable climates likely maintain diversity in the north while restriction to high elevations maintains relatively low diversity despite greater stability in southern regions. Results suggest genetic diversity gradients reflect long-term differences in N<sub>e</sub> dynamics and also emphasize the unique effects of isolation on insular habitats for bumble bees. Patterns are discussed in the context of conservation under climate change.</p>

opencc-zeroJan 2023View details →
dryad40/100

Population demographic history and evolutionary rescue: Influence of a bottleneck event

<p class="p1">Rapid environmental change presents a significant challenge to the persistence of natural populations. Rapid adaptation that increases population growth, enabling populations that declined following severe environmental change to grow and avoid extinction, is called evolutionary rescue. Numerous studies have shown that evolutionary rescue can indeed prevent extinction. Here, we extend those results by considering the demographic history of populations. To evaluate how demographic history influences evolutionary rescue, we created 80 populations of red flour beetle, <em>Tribolium castaneum</em>, with three classes of demographic history: diverse populations that did not experience a bottleneck, and populations that experienced either an intermediate or a strong bottleneck. We subjected these populations to a new and challenging environment for six discrete generations and tracked extinction and population size. Populations that did not experience a bottleneck in their demographic history avoided extinction entirely, while more than 20% of populations that experienced an intermediate or strong bottleneck went extinct. Similarly, among the extant populations at the end of the experiment, adaptation increased the growth rate in the novel environment the most for populations that had not experienced a bottleneck in their history. Taken together, these results highlight the importance of considering the demographic history of populations to make useful and effective conservation decisions and management strategies for populations experiencing environmental change that pushes them toward extinction.</p>

opencc-zeroJul 2023View details →
dryad40/100

Population genomics reveals demographic history and climate adaptation in Japanese Arabidopsis halleri

Open the record for dataset details and reuse information.

publicSep 2024View details →
dryad40/100

Data from: Evolutionary and demographic history of the Californian scrub white oak species complex: an integrative approach

Open the record for dataset details and reuse information.

publicNov 2015View details →
dryad40/100

Phylogenomics, introgression, and demographic history of South American true toads (Rhinella)

Open the record for dataset details and reuse information.

publicJan 2023View details →
dryad40/100

Joint analysis of microsatellites and flanking sequences enlightens complex demographic history of interspecific gene flow and vicariance in rear-edge oak populations

Open the record for dataset details and reuse information.

publicJun 2022View details →
dryad40/100

Whole genome demographic models indicate divergent effective population size histories shape contemporary genetic diversity gradients in a montane bumble bee

Open the record for dataset details and reuse information.

publicJan 2023View details →
dryad40/100

Data from: Revised age estimates for Northern Resident killer whales (Orcinus orca) based on observed life-history events and demographic discounting

Open the record for dataset details and reuse information.

publicFeb 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record