Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
19,487
datasets available to search
ShareScore release 0.7.1
Dataset results
19,487 results for “Populations”
Variation and overlaps in non-breeding regions of three less-studied breeding populations of the Eurasian Reed Warbler: Figures and datasheet
<p>Non-breeding regions of Eurasian Reed Warbler, <em>Acrocephalus scirpaceus</em> breeding in Finland, Jordan, and Kazakhstan as predicted with hydrogen stable isotope analysis.</p> <p>file name corresponds to the following figure titles:</p> <p>"Fig 2": <strong><span>Figure 2:</span></strong><span> The predicted non-breeding regions of Eurasian Reed Warblers breeding in Finland, Jordan, and Kazakhstan.</span></p> <p><span><span>"Fig 3": <strong><span>Figure </span></strong><span>3: Predicted non-breeding regions for Eurasian Reed Warblers sampled at a breeding site in Finland assigned to isotopic clusters A and </span><span><span>B.</span></span></span></span></p> <p><span><span></span></span><span></span><span><span><span>"Fig 4": </span></span></span><strong><span>Figure 4:</span></strong><span> Predicted within population average non-breeding regions for Eurasian Reed Warblers assigned to clusters A-F breeding in Jordan </span><span>(see Table S1 for details).</span></p> <p><span>Also uploaded is a csv file containing the isotopic signature of the birds. The file name is "data_Acrocephalus_scirpaceus"</span></p>
Two decades of body length measurements in size-structured larval and juvenile fish populations in English rivers.
<p>Long term ecological datasets are valuable in providing context and understanding to complex ecological processes that occur over broad temporal scales, and provide a baseline for analysing change. Monitoring of fish populations in UK waterbodies and elsewhere is typically through measuring the length of individual fish caught in surveys. Through this method, the age structure of fish populations can be determined, as well as over winer survival rates and future recruitment success and cohort sizes can be predicted. The larval and juvenile period are when fish are considered most vulnerable to predation, competition, disease and environmental perturbations. </p> <p><br>This study presents the first long-term larval and juvenile fish lengths dataset for 67 survey sites over two decades (1999-2018) from the rivers Ancholme, Warwickshire Avon, Don, Trent, and Yorkshire Ouse (including the Swale, Ure, Nidd and Wharfe) in the United Kingdom. These rivers represent a range of topographical and biotopical characteristics. For the majority of this study, surveys were conducted on a monthly or fortnightly basis making both annual and seasonal analyses of size structure, growth and body length possible. Although there is some variation in the sampling frequency and some locations varied throughout the study according to requirements. In total, more than 380,000 larval or juvenile fish of 30 species were measured, likely representing one of the most comprehensive datasets of its type.</p> <p>Surveys were conducted in river margins, where the velocity was slowest and larval and juvenile fish tend to aggregate. Fish were captured using a 25 x 3 m micromesh (3 mm mesh size) seine net that was set in a rectangle parallel to the bank. This net capture fish as small as 5 mm and is the most appropriate method of catching larvae and juvenile fish, although occasionally some larger adult fish may have also been captured and measured as part of this dataset for completeness. All fish were identified to species and measured to standard length (mm) and released at the point of capture. The exception was the smallest larvae, which were euthanised with an overdose of methanesulphonate (MS-222) and preserved in 4% formalin solution for microscopic examination.</p> <p><br>The dataset contains 384,090 rows and 13 columns. Each row corresponds to a single fish that was measured at each site and date. Associated site information (site name, location, area fished (m<sup>2</sup>) and survey date) is reported for each row. When only a fraction of the catch was processed, the sub-sample size was reflected in the Count column (e.g. when half the sample was processed, the numbers of fish measured or only counted were multiplied by two). This enables accurate densities to be calculated as the total number of both measured and unmeasured fish is recorded.</p> <p>Description of columns found in the dataset:</p> <p> </p> <table> <tbody> <tr> <td> <p><strong>Column heading</strong></p> </td> <td> <p><strong>Column description</strong></p> </td> <td> <p><strong>Data type</strong></p> </td> <td> <p><strong>Units</strong></p> </td> </tr> <tr> <td> <p>Fish _Catchment</p> </td> <td> <p>The river catchment/basin location of each fish site</p> </td> <td> <p>Text</p> </td> <td> <p>n/a</p> </td> </tr> <tr> <td> <p>Fish_River</p> </td> <td> <p>The river/watercourse location of each fish site.</p> </td> <td> <p>Text</p> </td> <td> <p>n/a</p> </td> </tr> <tr> <td> <p>Fish_SiteName</p> </td> <td> <p>The name of each fish site</p> </td> <td> <p>Text</p> </td> <td> <p>n/a</p> </td> </tr> <tr> <td> <p>Fish_Latitude</p> </td> <td> <p>The latitude of each fish site (WGS 1984)</p> </td> <td> <p>Integer</p> </td> <td> <p>Decimal degrees</p> </td> </tr> <tr> <td> <p>Fish_Longitude</p> </td> <td> <p>The longitude of each fish site (WGS 1984)</p> </td> <td> <p>Integer</p> </td> <td> <p>Decimal degrees</p> </td> </tr> <tr> <td> <p>Fish_Area</p> </td> <td> <p>Area of fish site surveyed</p> </td> <td> <p>Integer</p> </td> <td> <p>m<sup>-2</sup></p> </td> </tr> <tr> <td> <p>Fish_SurveyDate</p> </td> <td> <p>Date fish survey was carried out</p> </td> <td> <p>Integer</p> </td> <td> <p>dd/mm/yyyy</p> </td> </tr> <tr> <td> <p>Fish_Year</p> </td> <td> <p>Year fish survey was carried out</p> </td> <td> <p>Integer</p> </td> <td> <p>yyyy</p> </td> </tr> <tr> <td> <p>Common_Name</p> </td> <td> <p>The common/vernacular name of each fish taxon recorded in the dataset.</p> </td> <td> <p>Text</p> </td> <td> <p>n/a</p> </td> </tr> <tr> <td> <p>Latin_Name</p> </td> <td> <p>The scientific name of each fish taxon recorded in the dataset</p> </td> <td> <p>Text</p> </td> <td> <p>n/a</p> </td> </tr> <tr> <td> <p>Net_Number</p> </td> <td> <p>The net number the fish in a given survey were caught on</p> </td> <td> <p>Integer</p> </td> <td> <p>n/a</p> </td> </tr> <tr> <td> <p>Length_mm</p> </td> <td> <p>Length of individual fish caught</p> </td> <td> <p>Integer</p> </td> <td> <p>mm</p> </td> </tr> <tr> <td> <p>Count</p> </td> <td> <p>Count of fish caught accounting for sub- sampling</p> </td> <td> <p>Integer</p> </td> <td> <p>Number of fish</p> </td> </tr> </tbody> </table> <p> </p>
COVID19 Flow-Maps Population data
<p><strong>Daily population and trips per person data from Spain 2020-2021</strong></p> <p>This repository contains daily population records based on a study conducted by the MITMA, that analysed the mobility and distribution of the population in Spain from February 14th 2020 to May 9th 2021. The study is based on a sample of more than 13 million anonymised mobile phone lines provided by a single mobile operator whose subscribers are evenly distributed.</p> <p>For more information on the data visit: <a href="https://www.mitma.gob.es/ministerio/covid-19/evolucion-movilidad-big-data">https://www.mitma.gob.es/ministerio/covid-19/evolucion-movilidad-big-data</a></p> <p>Data provided by MITMA is related to the layer mitma_mov. For the rest of the layers, the population was estimated using the population grid from GEOSTAT: <a href="https://ec.europa.eu/eurostat/web/gisco/geodata/reference-data/population-distribution-demography/geostat">https://ec.europa.eu/eurostat/web/gisco/geodata/reference-data/population-distribution-demography/geostat</a></p>
Data and software supporting the manuscript 'The population frequency of human mitochondrial DNA variants is highly dependent upon mutational bias'
<p>Next-generation sequencing can quickly reveal genetic variation potentially linked to heritable disease. As databases encompassing human variation continue to expand, rare variants have been of high interest, since the frequency of a variant is expected to be low if the genetic change leads to a loss of fitness or fecundity. However, the use of variant frequency when seeking genomic changes linked to disease remains very challenging. Here, we explore the role of selection in controlling human variant frequency using the HelixMT database, which encompasses hundreds of thousands of mitochondrial DNA (mtDNA) samples. We find that a substantial number of synonymous substitutions, which have no effect on protein sequence, were never encountered in this large study, while many other synonymous changes are found at very low frequencies. Further analyses of human and mammalian mtDNA datasets indicate that the population frequency of synonymous variants is predominantly determined by mutational biases rather than by strong selection acting upon nucleotide choice. Our work has important implications that extend to the interpretation of variant frequency for non-synonymous substitutions. </p> <p> </p>
Data from: Complex population structure and haplotype patterns in Western Europe honey bee from sequencing a large panel of haploid drones
<p>This vcf file contains 7.023.689 SNPs and 870 honey bee samples, as described in the paper "Complex population structure and haplotype patterns in Western Europe honey bee from sequencing a large panel of haploid drones" by Wragg et al., available at https://doi.org/10.1101/2021.09.20.460798 as preprint.</p> <p>Eight hundred and seventy haploid drone samples from several honey bee subspecies hybrids were sequenced and aligned to the HAv3.1 reference genome. Sequence read alignment and genotyping quality filters were used to obtain a selection of 7.023.689 high-quality SNPs. The file Diversity_Study_629_Samples.txt corresponds to the 629 unique samples that were used for the diversity study described in the paper and can be used to recreate the restricted diversity dataset using bcftools or an equivalent software.</p> <p>Having sequenced haploid drones, heterozygous SNPs resulting from duplicated regions could be filtered out and the data is phased.</p>
The Thousand-Pulsar-Array program on MeerKAT -- IX. The time-averaged properties of the observed pulsar population: data set
<p>This archive contains pulsar data presented as part of the MNRAS paper: <em>"The Thousand-Pulsar-Array program on MeerKAT -- IX. The time-averaged properties of the observed pulsar population"</em>.</p> <p>Folded, time-averaged pulse profiles (4 Stokes parameters, 8 frequency channels, 1024 time bins across the period) of the 1271 pulsars listed in Table 1 of the MNRAS paper are included in the ar_files.zip. Ephemerides of these pulsars (as used in the MNRAS paper) are included in the eph_files.zip. The pulsar data are readable by the PSRCHIVE package, see e.g. van Straten et al., Astronomical Research and Technology 9, 237 (2012).</p> <p>Tables 1, 5, and 6 from the MNRAS paper are included in tables_files.zip as .csv files. The file column_descriptions.txt describes the quantities in columns of these tables.<br> </p>
Data from: Flock size and structure influence reproductive success in four species of flamingo in 540 captive populations worldwide
<p><strong>Summary</strong></p> <p>This dataset accompanies the publication "<strong>Flock size and structure influence reproductive success in four species of flamingo in 540 captive populations worldwide</strong>" published in Zoo Biology. It contains anonymised data from 540 captive flamingo populations, and includes the four species: <em>Phoeniconaias minor, Phoenicopterus chilensis, Phoenicopterus roseus</em> and<em> Phoenicopterus ruber</em>. Data were sourced from the Zoological Information Management System (ZIMS), operated by Species360 (https://www.species360.org/). ZIMS is the largest real-time database of comprehensive and standardized information spanning more than 1,200 zoological collections globally, and provides the number of institutions currently managing each flamingo species and both their current and historic population sizes. These data were used to investigate the relationship between reproductive success and both flock size, and structure, on a global scale.</p> <p>This dataset also contains climatic data provided by WorldClim, which were used to assess the influence of climatic variables on captive flamingo reproductive success globally. The WorldClim database averages 19 different climatic variables derived from monthly temperature and rainfall values at a 1 km spatial resolution for the period 1970-2000. Using geographic coordinates (latitude and longitude) we calculated several climatic metrics for each institution. </p> <p> </p> <p><strong>Description of the Dataset</strong></p> <p>One file is provided for each species (<em>P. minor, P. chilensis, P. roseus </em>and <em>P. ruber</em>) as a csv file. Each file contains the following 15 columns:</p> <ul> <li><strong>Institution Code: </strong>An anonymous code used to identify individual zoological institutions. </li> <li><strong>Country: </strong>The country where the institution is located.</li> <li><strong>Year: </strong>Current year (<em>t</em>).</li> <li><strong>Flock Size:</strong> Flock size in year <em>t.</em></li> <li><strong>Males: </strong>The number of males in the flock in year <em>t.</em> </li> <li><strong>Females:</strong> The number of females in the flock in year <em>t.</em></li> <li><strong>Unsexed:</strong> The number of unsexed individuals in the flock in year <em>t.</em></li> <li><strong>Proportion of Females: </strong>The proportion of the flock made up of female individuals in year <em>t</em>. </li> <li><strong>Proportion of Unsexed:</strong> The proportion of the flock made up of unsexed individuals in year <em>t.</em></li> <li><strong>Hatches:</strong> Number of birds hatched in year <em>t.</em></li> <li><strong>Proportion of Additions:</strong> The proportion of the flock in year <em>t</em> made up of additions from year <em>t-1</em> (not including new birds hatched into the flock).</li> <li><strong>MAP: </strong>Mean annual precipitation (mm).</li> <li><strong>MAT: </strong>Mean annual temperature (°C).</li> <li><strong>MAP Var: </strong>Mean annual variation in precipitation (MAP coefficient of variation).</li> <li><strong>MAT Var: </strong>Mean annual variation in temperature (MAT standard deviation).</li> </ul> <p>Note: Mean Annual Temperature (MAT) is provided by WorldClim as °C multiplied by 10, and similarly mean annual variation in temperature as MAT standard deviation multiplied by 100. In the corresponding publication, both were divided (by 10 and 100 respectively) prior to modelling to avoid confusion in the units used.</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We acknowledge and thank all Species360 member institutions for their continued support and data input. The research which data refers to was funded by the Irish Research Council Laureate Awards 2017/2018 IRCLA/2017/60 to Y.M.B. Additionally, S.Q.S. received funding from the International Max Planck Research School for Organismal Biology. The Species360 Conservation Science Alliance would like to thank their sponsors: the World Association of Zoos and Aquariums, Wildlife Reserves of Singapore, and Copenhagen Zoo. </p> <p> </p> <p><strong>Disclaimer</strong></p> <p>Despite our best efforts at screening the data for errors and inconsistencies, some information could be erroneous. Similarly, data contained within ZIMS are based on submitted records from individual institutions, and are not subject to editorial verification, potentially permitting errors or failure to update species holdings etc. Despite this, ZIMS represents the only global database of zoo collection composition records, and as a result, is used by the IUCN, Convention on International Trade in Endangered Species (CITES), the Wildlife Trade Monitoring Network (TRAFFIC), United States Fish and Wildlife Service (USFWS) and Department for Environment, Food and Rural Affairs (DEFRA). </p> <p> </p> <p><strong>Credit</strong></p> <p>If you use this dataset, please cite the corresponding publication:</p> <p>Mooney, A., Teare, J. A., Staerk, J.,Smeele, S. Q., Rose, P., Edell, R. H., King, C. E., Conrad, L., & Buckley, Y. M. (2023). Flock size and structure influence reproductive success in four species of flamingo in 540 captive populations worldwide.<em> Zoo Biology</em>, 1–14. <a href="https://doi.org/10.1002/zoo.21753">https://doi.org/10.1002/zoo.21753</a></p> <p> </p> <p> </p>
Relative density variations of common vole population based on index transect, Septfontaines - Le Souillot, France (1990-2000)
<p>Transects were walked from village to village along a transect line. Common vole (<em>Microtus arvalis</em>) activity indices were recorded in every ten pace interval from October 1990 to April 2000. In 2014, the geographical coordinates of each interval has been computed by spatial interpolation based on georeferenced maps. Therefore, users must be aware that individual locations of intervals are unprecise, but not the general bearing of the transect in the landscape and interval succession. See articles published for reference and more details.</p> <p>During the same time span, small mammmals (including common voles) were sampled using live-trapping, see <a href="https://doi.org/10.5281/zenodo.6997316">10.5281/zenodo.6997316</a></p> <p><strong>FILE DESCRIPTION:</strong></p> <p><a href="https://zenodo.org/record/7544358/files/db.txt?download=1">db.txt </a>index transect file</p> <ul> <li>name: transect name</li> <li>date: on eight digits, '19921014' reads 14/10/1992</li> <li>ID: interval ID = number (within a given transect at a given date)</li> <li>Habitat: (indicative) the habitat category crossed. Just mentioned when passing from one category to the other; the following intervals are assumed to belong to this habitat</li> <li>ma1: number of <em>Microtus</em> holes; A, 1-5 holes; B, 6-10 holes; C > 10 holes</li> <li>ma2: answered only if A, B, or C are defined in ma1; NA, not answered (ma1 not defined), 0, zero faeces, 1 some faeces or fresh indices (runways with grass freshly cut, etc.); 2 many faeces in heaps</li> <li>long: longitude (WGS84)</li> <li>lat: latitude (WGS84)</li> </ul> <p><a href="https://zenodo.org/record/7544358/files/StudyAreaBoundingBox.kml?download=1">StudyAreaBoundingBox.kml</a> Bounding box of the study area.</p>
Simulation of the Galactic field millisecond pulsar population and its gamma- and X-ray emission
<p>Monte Carlo simulation of the millisecond pulsar population in the Galactic field. The simulation includes four spatial components:</p> <ul> <li>the disk;</li> <li>the boxy bulge;</li> <li>the nuclear stellar cluster;</li> <li>the nuclear stellar disk.</li> </ul> <p>The last 3 components together form the Galactic bulge. There is one file per component, each containing at least 100 Monte Carlo simulations. Each line contains:</p> <ul> <li>the longitude L in deg;</li> <li>the latitude B in deg;</li> <li>the line of sight S in kpc;</li> <li>the 0.1-100 GeV gamma-ray flux in erg/cm^2/s;</li> <li>the X-ray spectral index;</li> <li>the gamma-to-X flux ratio, where the gamma-ray flux is the same as in the fourth column and the X-ray flux is the 2-10 keV unabsorbed one</li> </ul> <pre>of a simulated MSP. More information about the simulation can be found in the related paper. </pre>
Population genomics reveals differences in genetic structure between two endemic arboreal rodent species in threatened cloud forest habitat
<p>SNPs obtained by UNEAK pipeline for <em>Habromys schmidlyi </em>and <em>Reithrodontomys microdon</em>. </p> <p>Pleae cite as: </p> <p>Colunga-Salas P., T Marines-Macías, G Hernández-Canchola, S Barbosa, C Ramírez, JB Searle, L León-Paniagua. 2022. <strong>Population genomics reveals differences in genetic structure between two endemic arboreal rodent species in threatened cloud forest habitat</strong>. Mammalian Reasearch. Doi: 10.1007/s13364-022-00667-x</p>
Netanyahu's and Abbas' speeches at the UNGA 2010-19, coded using a populism framework
<p>Databaset with speeches of Benjamin Netanyahu and Mahmoud Abbas before the United Nations General Assembly (UNGA) (2010-2019) coded using MAXQDA following populism multidimensional comparative framework by Olivas Osuna (2021).</p> <p>In this database syntactic units —sentences—are individually coded whenever they match the criteria corresponding to any of populism/anti-populism, re-bordering/de-bordering, religion and securitisation codes defined previously. See Olivas Osuna and Rama (2021) and Olivas Osuna (2022) for reference to the methodology and previous empirical applications.</p>
SEVN parameter file from the paper "Binary neutron star populations in the Milky Way" by Sgalletta et al., 2023
<p>The repository contains the runtime parameters used in the SEVN simulations analysed in the paper "Binary neutron star populations in the Milky Way" by Sgalletta et al., 2023.</p> <p><strong>Repository content: </strong></p> <p>- <em>used_params_Sgalletta2023.txt<br> </em>The file contains all the runtime parameters used in the SEVN simulations. The parameters that have been varied in different runs are indicated with **** and the explored values are reported in the comment. See the SEVN userguide (<a href="https://gitlab.com/sevncodes/sevn/-/blob/SEVN/resources/SEVN_userguide.pdf">https://gitlab.com/sevncodes/sevn/-/blob/SEVN/resources/SEVN_userguide.pdf</a>) for the description of each parameter </p> <p> </p>
The short gamma-ray burst population in a quasi-universal jet scenario: MCMC chains
<p>The paper "The short gamma-ray burst population in a quasi-universal jet scenario" (https://arxiv.org/abs/2306.15488) described an effort in modelling the short gamma-ray burst population under the assumption that all jets share the same angular profile.</p> <p>This repository contains <strong>emcee </strong>hdf5 files with the MCMC chains corresponding to the "full sample" and "flux-limited sample" analyses described in the paper.</p>
QuantMig microsimulation population projection model and migration scenarios for 31 European countries
<p>This open data deposit contains the data and model code of QuantMig-Mic microsimulation population projection model for 31 European countries and accompanies deliverables D8.3: Model outputs for dissemination and D8.1: Microsimulation projection model.</p> <p>This Zenodo deposit contains datasets of model input (baseline population and immigration database) and output data (demography and components output tables) and model code with parameters of the Baseline scenario (termed Default in the model code) deposited in QuantMig_mic.zip file. To view the code the users must first install MODGEN software (or can view code files in Visual Studio). All scenarios share the same parameters except the immigrant population - to change the immigration assumptions the users can import immigration assumptions for any other scenario from the ImmigDataBase.csv and change it in the immigration module using the MODGEN user interface or using Visual Studio.</p> <p><strong>The file structure and codebook for the data files is included in the cover note file "readme_quantmig_datasets.pdf"</strong></p> <p>Detailed <strong>information about the QuantMig-Mic microsimulation model, its modules and parameters</strong>:</p> <p>Marois, G., Potančoková, M., González-Leonardo, M. (2023) QuantMig-Mic microsimulation population projection model. QuantMig Project Deliverable D8.2. International Institute for Applied Systems Analysis (IIASA), Austria. Available at:http://quantmig.eu/res/files/QuantMig_IIASA_Deliverable%20D8.2_forsubmission_21-07-2023_corr.pdf</p> <p>Instructions how to install MODGEN can be found in:</p> <p>Marois, G. and Potančoková, M. (2022) QuantMig-mic microsimulation tool. QuantMig Project Deliverable D8.1. International Institute for Applied Systems Analysis (IIASA). http://quantmig.eu/res/files/QuantMig_IIASA_Deliverable%20D8.1%20v1.1.pdf </p> <p>Detailed <strong>information about the QuantMig migration scenarios</strong> can be found in:</p> <p>Marois, G., Potančoková, M., González-Leonardo, M. (2023) QuantMig-Mic microsimulation population projection model. QuantMig Project Deliverable D8.2. International Institute for Applied Systems Analysis (IIASA), Austria. Available at:http://quantmig.eu/res/files/QuantMig_IIASA_Deliverable%20D8.2_forsubmission_21-07-2023_corr.pdf</p> <p>A <strong>guide to the datasets</strong> and the codebook can be found in: <strong>readme_quantmig_datasets.pdf</strong></p> <p><strong>Countries included in the model: </strong></p> <p>Austria, Belgium, Bulgaria, Croatia, Czechia, Cyprus, Denmark, Estonia, Finland, France, Germany, Greece, Hungary, Iceland, Ireland, Italy, Latvia, Lithuania, Luxembourg, Malta, Netherlands, Norway, Poland, Portugal, Romania, Slovakia, Slovenia, Spain, Sweden, Switzerland, United Kingdom</p> <p> </p>
Micro elemental composition of Pike-perch (Sander lucioperca) population in Lipno Reservoir, Czechia
<p>This dataset contains the information on the micro elemental composition of Sagitta otoliths of Pike-Perch (<i>Sander lucioperca</i>) collected in Lipno Reservoir (Czechia). The dataset covers a wide range of micro elemental components (barium, calcium, copper, potassium, lithium, magnesium, manganese, sodium, rubidium, strontium and zinc) obtained from the otolith cores and rims of these fish specimens. The dataset includes readings from Pike-Perch directly collected in Lipno Reservoir, as well as from those reared in facilities and later introduced into the reservoir.</p>
Inferring whole-genome histories in large population datasets: inferred tree sequences for 1000 Genomes
<p>Tree sequences inferred for the 1000 Genomes phase 3 autosomes using <a href="https://tsinfer.readthedocs.io/">tsinfer</a> version 0.1.4 and compressed using <a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. Tree sequences can be decompressed as follows:</p> <pre><code class="language-bash">$ tsunzip 1kg_chr1.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed using <a href="https://tskit.readthedocs.io">tskit</a>. </p> <pre><code class="language-python">import tskit ts = tskit.load("1kg_chr1.trees") # ts is an instance of tskit.TreeSequence print("Chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with individuals and populations was derived from the original <a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/working/20130606_sample_info/20130606_g1k.ped">source</a> and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code class="language-python">import tskit import json ts = tskit.load("1kg_chr1.trees") ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code class="language-python">pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on <a href="https://github.com/mcveanlab/treeseq-inference/tree/master/human-data">GitHub</a>.</p>
Inferring whole-genome histories in large population datasets: inferred tree sequences for Simons Genome Diversity Project
<p>Tree sequences inferred for the SGDP autosomes using <a href="https://tsinfer.readthedocs.io/">tsinfer</a> version 0.1.4 and compressed using <a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. Tree sequences can be decompressed as follows:</p> <pre><code class="language-bash">$ tsunzip sgdp_chr1.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed using <a href="https://tskit.readthedocs.io">tskit</a>. </p> <pre><code class="language-python">import tskit ts = tskit.load("sgdp_chr1.trees") # ts is an instance of tskit.TreeSequence print("Chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with individuals and populations was derived from the original <a href="https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/SGDP_metadata.279public.21signedLetter.samples.txt">source</a> and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code class="language-python">import tskit import json ts = tskit.load("sgdp_chr1.trees") ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code class="language-python">pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on <a href="https://github.com/mcveanlab/treeseq-inference/tree/master/human-data">GitHub</a>.</p>
Central Valley Project, Genetic Determination of Population of Origin 2011-2021
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Knights Landing, California Department of Fish and Wildlife, Genetic Determination of Population of Origin 2017 through 2019
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Sacramento trawl, Delta Juvenile Fish Monitoring Program, Genetic Determination of Population of Origin 2017-2021
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.