Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,940
datasets available to search
ShareScore release 0.9.0
Dataset results
1,940 results for “data sample”
Global ECMWF Fire Forecasting system - sample data for wildfires in Siberia on 23-25 July 2019
<p>The European Centre for Medium-Range Weather Forecasts (<a href="https://www.ecmwf.int/">ECMWF</a>) produces daily fire danger forecasts and reanalysis products from the Global ECMWF Fire Forecast (<a href="https://git.ecmwf.int//projects/CEMSF/repos/geff/browse">GEFF</a>) model. Reanalysis is available through the Copernicus Climate Data Store (<a href="https://cds.climate.copernicus.eu/cdsapp#%21/dataset/cems-fire-historical">CDS</a>) while the medium-range real-time forecast is available through the <a href="https://effis.jrc.ec.europa.eu/static/effis_current_situation/public/index.html">EFFIS</a> and <a href="https://gwis.jrc.ec.europa.eu/static/gwis_current_situation/public/index.html">GWIS</a> platforms.</p> <p>This repository provides FWI sample datasets for the assessment of the Siberian wildfires occurred on 23-25 July 2019:</p> <ul> <li> <p>GEFF-reanalysis, which provides historical records of fire danger conditions</p> <ul> <li> <p>e5_hr, this folder contains deterministic model outputs</p> </li> <li> <p>e5_en, this folder contains probabilistic model outputs (made of 10 ensemble members)</p> </li> </ul> </li> <li> <p>GEFF-realtime provides real-time forecasts generated using weather forcings from the model cycle 46r1 of the ECMWF’s Integrated Forecasting System (IFS).</p> <ul> <li> <p>rt_hr, this folder contains high-resolution deterministic forecasts (~9 Km)</p> </li> <li> <p>rt_en, this folder contains probabilistic forecasts (~18Km)</p> </li> </ul> </li> <li> <p>Geographical bounding box: lon_min = 40, lon_max = 180, lat_min = -10, lat_max = 90</p> </li> </ul> <p><strong>Please note, the sample data provided in this repository is intended to be used for education purposes only (e.g. training courses).</strong></p> <p>These products have been developed as part of the EU-funded Copernicus Emergency Management Services (<a href="https://emergency.copernicus.eu/">CEMS</a>) and complement other Copernicus products related to fire, such as the biomass-burning emissions made available by the Copernicus Atmosphere Monitoring Service (<a href="https://atmosphere.copernicus.eu/">CAMS</a>). The development of the GEFF modelling system was funded through a third-party agreement with the European Commission’s Joint Research Centre (<a href="https://ec.europa.eu/info/departments/joint-research-centre_en">JRC</a>). </p> <p>GEFF produces fire danger indices based on the Canadian Fire Weather index as well as the US and Australian fire danger models. GEFF datasets are under the Copernicus license, which provides users with free, full and open access to environmental data.</p> <p>For more information, please refer to the documentation on the <a href="http://datastore.copernicus-climate.eu/c3s/published-forms/c3sprod/cems-fire-historical/Fire_In_CDS.pdf">CDS</a> and on the <a href="https://effis.jrc.ec.europa.eu/about-effis/technical-background/fire-danger-forecast/">EFFIS website</a>.</p>
Rt-Cloud rtAttenPenn Sample DICOM data
<p>This upload contains the same data as published in <a href="https://doi.org/10.5281/zenodo.3677090">our previous zenodo dataset upload</a>. Unlike our previous upload, this version contains data after transferring the DICOMs directly from the Siemens Skyra 3T to our Linux machine (as done in real-time experiments). The purpose of this separate upload is to serve as sample data for our <a href="https://github.com/brainiak/rt-cloud">real-time cloud software</a>, for the r<a href="https://github.com/brainiak/rtAttenPenn_cloud">tAttenPenn project</a>. A previous version of this same data was uploaded for a different rt-cloud <a href="https://github.com/amennen/amygActivation">sample project</a> <a href="https://doi.org/10.5281/zenodo.3862783">(zenodo link here)</a>. For the current sample project, we are additionally including a T1w anatomical scan for real-time registration. All brain data are contributed by author S.A.N. and are authorized for non-anonymized distribution.</p>
Experimental Data Set for the study "Exploratory Landscape Analysis is Strongly Sensitive to the Sampling Strategy"
<p>This are the feature values used in the study "Exploratory Landscape Analysis is Strongly Sensitive to the Sampling Strategy".</p> <p>The dataset regroups feature values for every "cheap" features available in the R package <em>flacco </em>and are computed using 5 sampling strategies and in dimension <span class="math-tex">\($d=5$\)</span>:</p> <ol> <li>Random: the classical Mersenne-Twister algorithm;</li> <li>Randu: a random number generator that is notoriously bad;</li> <li>LHS: a centered Latin Hypercube Design;</li> <li>iLHS: an improved Latin Hypercube Design;</li> <li>Sobol: points extracted from a Sobol' low-discrepancy sequence.</li> </ol> <p>The csv file <em>features_summury_dim_5_ppsn.csv </em>regroups 100 values for every features whereas <em>features_summury_dim_5_ppsn_median.csv </em>regroups for every feature the median of the 100 values.</p> <p>In the folder <em>PPSN_feature_plots</em> are the histograms of feature values on the 24 COCO functions for 3 sampling strategies: Random, LHS and Sobol.</p> <p>The Python file <em>sampling_ppsn.py</em> is the code used to generate the sample points from which the feature values are computed.</p> <p>The file <em>stats50_knn_dt.csv</em> provide the raw data of median and IQR (inter quartile interval) for the heatmaps and boxplots available in the paper.</p> <p>Finally, the files <em>results_classif_knn100.csv</em> (resp. dt) provide the accuracy of 100 classifications for every settings.</p> <p> </p>
Data for manuscript "The proportion of non-depressed subjects in a study sample strongly affects the results of psychometric analyses of depression symptoms"
<p>Data sets underlying the manuscript "The proportion of non-depressed subjects in a study sample strongly affects the results of psychometric analyses of depression symptoms" (doi: 10.1371/journal.pone.0235272). The file "data_table1.rds" contains the data underlying table 1 of the manuscript, whereas the file "data_figure1.rds" contains the data underlying figure 1. Both files are stored in RDS-format, which is a data format used in the programming language R. RDS-files can be loaded using R with the function readRDS(). Note that "data_figure1.rds" is a list of 22 elements. Elements 1, 21, and 22 are original data sets (depressed subjects, general population, and nondepressed subjects, respectively), while the remaining elements are lists each containing 500 newly generated data sets as described in the method section of the manuscript. For more details about the data please refer to the manuscript.</p>
Simulated data and results from "Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data"
<p>This dataset contains all the simulated data and the results of all the considered methods in the benchmark presented in "Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data" [Zaccaria & Raphael, 2018]. All the data in this dataset and the corresponding formats are fully described at <a href="https://github.com/raphael-group/hatchet-paper">https://github.com/raphael-group/hatchet-paper</a>. The folder <em>simulations</em> which contains the entire dataset has been compressed with standard <em>zip</em>.</p>
Data from: Sampling schemes and drift can bias admixture proportions inferred by STRUCTURE
<p><span>The interbreeding of individuals coming from genetically differentiated but incompletely isolated populations can lead to the formation of admixed populations, having important implications in ecology and evolution. In this simulation study, we evaluate how individual admixture proportions estimated by the software <span>structure</span> are quantitatively affected by different factors. Using various scenarios of admixture between two diverging populations, we found that unbalanced sampling from parental populations may seriously bias the inferred admixture proportions; moreover, proportionally large samples from the admixed population can also decrease the accuracy and precision of the inferences. As expected, weak differentiation between parental populations and drift after the admixture event strongly increase the biases caused by uneven sampling. We also show that admixture proportions are generally more biased when parental populations unequally contributed to the admixed population. Finally, with few exceptions, using a large number of markers reduces those biases, but using alternative priors for individual ancestry or the uncorrelated allele model only marginally affect the inference of admixture in most situations. We conclude that unbalanced sampling may cause important biases in the admixture proportions estimated by <span>structure</span>, especially when a small number of markers are used, and those biases can be worsened by the effect of drift and unequal genetic contribution of parental populations. Empirical studies should thus be careful with their sampling design and consider historical characteristics when using this software to estimate the ancestry of individuals from admixed populations.</span></p>
Soil properties and N-cycling community data for 324 samples
<p>Soil microbial communities, key players of many crucial ecosystem functions, are prone to degradation due to intensified environmental perturbations. However, few ecological concepts and practices have been developed for rescuing these degraded microbial communities. Here, we carried out a pot experiment (totally 324 pots) with or without arbuscular mycorrhizal fungi (AMF) inoculation and at three levels (one, three and six species) of plant diversity to disentangle how the AMF and vegetation rescue the soil nitrogen (N) -cycling microbial loop. Our results provide strong evidence that AMF inoculation improved the restoration of soil N-cycling microbial communities. This improved restoration is related to the role of AMF in enhancing interactions within the N-cycling microbial loop. Furthermore, increased plant diversity strengthened the role of AMF in rescuing degraded N-cycling microbial communities. Our findings provide novel insights into the roles of AMF and plant diversity in facilitating the rescue of microbial communities in degraded terrestrial ecosystems.</p>
Data from: Optimising sample sizes for animal distribution analysis using tracking data
<p><span>1. Knowledge of the spatial distribution of populations is fundamental to management plans for any species. When tracking data are used to describe distributions, it is sometimes assumed that the reported locations of individuals delineate the spatial extent of areas used by the target population.</span></p> <p><span>2. Here, we examine existing approaches to validate this assumption, highlight caveats, and propose a new method for a more informative assessment of the number of tracked animals (i.e. sample size) necessary to identify distribution patterns. We show how this assessment can be achieved by considering the heterogeneous use of habitats by a target species using the probabilistic property of a utilisation distribution. Our methods are compiled in the R package <i>SDLfilter</i>.</span></p> <p><span>3. We illustrate and compare the protocols underlying existing and new methods using conceptual models and demonstrate an application of our approach using a large satellite tracking data-set of flatback turtles, <i>Natator depressus, </i>tagged with accurate Fastloc-GPS tags (n = 69).</span></p> <p><span>4. Our approach has applicability for the post-hoc validation of sample sizes required for the robust estimation of distribution patterns across a wide range of taxa, populations and life history stages of animals.</span></p>
Figure 6 from: Bonello G, Grillo M, Cecchetto M, Giallain M, Granata A, Guglielmo L, Pane L, Schiaparelli S (2020) Distributional records of Ross Sea (Antarctica) planktic Copepoda from bibliographic data and samples curated at the Italian National Antarctic Museum (MNA): checklist of species collected in the Ross Sea sector from 1987 to 1995. ZooKeys 969: 1-22. https://doi.org/10.3897/zookeys.969.52334
Figure 6 Metridia gerlachei (Copepoda, Calanoida; female, MNA-12439) acquired with fluorescence microscopy (Congo Red, 1.5 mg/ml). This species is one of the most adapted species in the Antarctic region and can perform diel vertical migrations that highly influence the surrounding waters in terms of trophic relationships in Terra Nova Bay.
Figure 1 from: Bonello G, Grillo M, Cecchetto M, Giallain M, Granata A, Guglielmo L, Pane L, Schiaparelli S (2020) Distributional records of Ross Sea (Antarctica) planktic Copepoda from bibliographic data and samples curated at the Italian National Antarctic Museum (MNA): checklist of species collected in the Ross Sea sector from 1987 to 1995. ZooKeys 969: 1-22. https://doi.org/10.3897/zookeys.969.52334
Figure 1 Sampling stations for IIIrd (yellow), Vth (blue), and Xth (red) expedition a overview of spatial extent in Antarctica b sampling stations in the Western Ross Sea c focus on Terra Nova Bay sampling stations. This map was produced using the collection of datasets "Quantarctica" (Matsuoka et al. 2017) and QGIS (QGIS Development Team 2020).
Figure 5 from: Bonello G, Grillo M, Cecchetto M, Giallain M, Granata A, Guglielmo L, Pane L, Schiaparelli S (2020) Distributional records of Ross Sea (Antarctica) planktic Copepoda from bibliographic data and samples curated at the Italian National Antarctic Museum (MNA): checklist of species collected in the Ross Sea sector from 1987 to 1995. ZooKeys 969: 1-22. https://doi.org/10.3897/zookeys.969.52334
Figure 5 Paraeuchaeta exigua (Copepoda, Calanoida; female, MNA-12333) acquired with scanning electron microscopy (SEM). This is one of the most common species in the coastal area of Terra Nova Bay. It plays a key role in the neritic trophic chain and highly contributes to the total mesozooplanktic biomass.
Speed-Accuracy Trade-Offs in Sample-Based Decisions - Data
<p>Datasets for all experiments described in the paper "Speed-Accuracy Trade-Offs in Sample-Based Decisions - Data" published in JEP:General in 2020.</p>
In-time particle simulation (Itpas) - sample data
<p>Itpas is a Lagrangian particle dispersion model, online coupled to the weather forecast model COSMO.</p> <p>This archive contains a sample data set of the Itpas output and the associated meteorological data fields. </p> <p>The model domain covers an area from eastern Germany to central Poland including the cities of Berlin and Poznan (lower left corner: 51°N 12°E, upper right corner: 54.4°N 19.3°E) and has a size of 167 x 150 grid cells with a spatial resolution of 2.8 km and 50 vertical levels.<br> The data is assigned to May 31, 2017 form 08 UTC to 23 UTC</p> <p>l.nc -> Atmospheric data produced by the COSMO-model</p> <p>part_t001200_p001.nc<br> part_t001260_p001.nc -> Trajectories emitted in the morning from 08:50 to 09:45 UTC</p> <p>part_t001380_p001.nc<br> part_t001440_p001.nc -> Trajectories emitted at noon from 11:40 to 12:30 UTC</p> <p>mean_traj.txt -> position of the mean trajectory [time rot_lon rot_lat z_sea z_surface]<br> stda_traj.txt -> standard deviation of the mean trajectory</p> <p>The data can be used as sample data for the following software:<br> https://github.com/mttfst/trajectory-plot<br> https://github.com/mttfst/trajectory-cross-section</p>
Sampling data, colony runfiles, VCFs, and Rscripts for: genomic determination of reproductive mode in facultatively parthenogenetic Opiliones
<p>Sexual reproduction may pose myriad short-term costs to females. Despite these costs, sexual reproduction is near ubiquitous. Facultative parthenogenesis is theorized to mitigate some of the costs of sex, as individuals can participate in occasional sex to limit costs while obtaining many benefits. However, most theoretical models assume sexual reproduction is fixed following mating, with no possibility of clutches of mixed reproductive ontogeny. Therefore, we asked: if coercive males are present at high frequency in a population of facultative parthenogens, will their clutches be solely sexually produced, or will there be evidence of sexually and asexually-produced offspring? How will their offspring production compare to conspecifics in low frequency male populations? We addressed our questions by collecting females and egg clutches of the facultatively parthenogenetic Opiliones species <i>Leiobunum manubriatum </i>and <i>L. globosum</i>. In <i>L. manubriatum</i>, females from populations with few males were not significantly more fecund than females from populations with higher male relative frequency, despite the potential release of the former from sexual conflict. We used three genotyping methods along with a custom set of DNA capture probes to reveal that offspring of <i>L. manubriatum </i>from these high male populations were primarily produced via asexual reproduction. This is surprising because sex ratios in these southern populations approach equality, increasing the probability for females to encounter mates and produce offspring sexually. We additionally found evidence for reproductive polymorphisms within populations. Rapid and accurate SNP genotyping data will continue to allow us to address broader evolutionary questions regarding the role of facultative reproductive modes in the maintenance of sex.</p>
Presence-absence sampling for estimating plant density using survey data with variable plot size
1. Presence-absence sampling is an important method for monitoring state and change of both individual plant species and communities. With this method only the presence or absence of the target species is recorded on plots and thus the method is straightforward to apply and less prone to surveyor judgment compared to other vegetation monitoring methods. However, in the basic setting all plots must be equally large or otherwise it is unclear how data should be analyzed. In this study we propose and evaluate five different methods for estimating plant density based on presence-absence registrations from surveys with variable plot sizes. 2. Using artificial plant population data as well as empirical data from the Swedish National Forest Inventory we evaluated the performance of the proposed methods. The main analysis was conducted through sampling simulation in the artificial populations, whereby bias and variance of density estimators for the different methods were quantified and compared. 3. Both for state and change estimation of plant density, we found that the best method to handle variable plot size was to perform generalized least squares regression, using plot size as an independent variable. Methods where plots smaller than a certain threshold were excluded or their registrations recalculated were, however, almost as good. Using all registrations as if they were obtained from plots with the nominal plot size resulted in substantial bias. 4. Our findings are important for plant population studies in a wide range of environmental monitoring programmes. In these programmes plots are typically randomly laid out and may be located across boundaries between different land use or land cover classes, resulting in subplots of variable size. Such splitting of plots is common when large plots are used, e.g. with the 100 m2 plots used in the Swedish National Forest Inventory. Our methods overcome problems to estimate plant density from presence-absence data observed in plots that vary in size.
UPLC-IMS-QTOF data from road runoff samples - mz5 files with scans in centroid spectrum format
<p>This dataset has been acquired from road runoff samples collected as part of the Roulepur project (Office Français de la Biodiversité, Agence de l’eau de Seine-Normandie).</p> <p>All samples were extracted by SPE with HLB cartridges at pH 2. Data was acquired with a UPLC-IMS-QTOF (Waters Vion) running Unifi, in HDMSe mode. Samples were injected in a random sequence (see order in samplemetadata.csv file). A list of markers was obtained from Unifi after alignment across all samples, and exported as a markertable.csv file. Each detected marker has m/z, retention time, drift time (ion mobility) values, and intensity values in all samples. Raw data was converted using Proteowizard MSConvert version 3.0.19014-f9d5b8a3b, using zlib compression, combining ion mobility scans, with the CWT peakpicking algorithm (snr=0, peakSpace=0) to transform profile to centroid data, and the zeroSamples option (removeExtra 1-). The chosen output format was .mz5.</p> <p>For file naming explanation, see the readme.txt file.</p>
Microsatellite genotype data from seven loci for a phylogeographic/population genetic study of the South African endemic freshwater crab Potamonautes lividus sampled from eight localities in the KwaZulu-Natal and Eastern Cape provinces in South Africa
<ol> <li>During the present study, the phylogeography of the only southern African IUCN Red Listed vulnerable (VU) freshwater crab, <i>Potamonautes lividus</i> was investigated by surveying several localities in the Eastern Cape and KwaZulu-Natal provinces in South Africa. Both nuclear and mitochondrial DNA markers were used, and it was hypothesized, that marked genetic differentiation should be present, while niche modeling was undertaken to explore the distribution of the species along the east coast of South Africa. <span><span>Further, the shortfalls in the present approach to IUCN Red Listing, as illustrated by a vulnerable species of crabs are discussed</span></span>.</li> <li>Results from the mtDNA revealed the presence of two haploclades confined to specimens from the two provinces respectively and the general absence of maternal dispersal; a fact that was further validated by the marked <i>F</i><sub>ST</sub> data and high F<sub>ST</sub>. Within the Eastern Cape haploclade, low frequency maternal dispersal occurred, corroborated by the low F<sub>ST</sub>. In contrast, no haplotypes were shared in the KwaZulu-Natal haploclade a fact corroborated by marked F<sub>ST</sub> differences. </li> <li>The microsatellite data demonstrated the presence of higher frequency, possibly paternally biased dispersal of specimens between the Eastern Cape and KwaZulu-Natal provinces. Our results suggest that presence of two distinct management units within <i>P. lividus</i>. Divergence time estimation suggest a late Pleistocene cladogenesis between the Eastern Cape and KwaZulu-Natal haploclades. </li> <li>Considering the presence of <i>P. lividus</i> in several newly collected nature conservation areas in both provinces, and its potential presence in the intermediary area based on the MAXENT niche modeling, our data suggest the species IUCN Red Listing status should be downgraded to LC.</li> <li>A comparison of all the EN, VU and CR IUCN Red Listed freshwater crabs for the entire Afrotropical region reveals the lack of recent sampling in the three biodiversity hotspots in West, Central and East Africa, with mountainous areas containing a disproportionate number of species with most species being devoid of phylogeographic study. </li> </ol>
Data from: The impact of the tree prior on molecular dating of data sets containing a mixture of inter- and intraspecies sampling
In Bayesian phylogenetic analyses of genetic data, prior probability distributions need to be specified for the model parameters, including the tree. When Bayesian methods are used for molecular dating, available tree priors include those designed for species-level data, such as the pure-birth and birth-death priors, and coalescent-based priors designed for population-level data. However, molecular dating methods are frequently applied to data sets that include multiple individuals across multiple species. Such data sets violate the assumptions of both the speciation and coalescent-based tree priors, making it unclear which should be chosen and whether this choice can affect the estimation of node times. To investigate this problem, we used a simulation approach to produce data sets with different proportions of within- and between-species sampling under the multispecies coalescent model. These data sets were then analysed under pure-birth, birth-death, constant-size coalescent, and skyline coalescent tree priors. We also explored the ability of Bayesian model testing to select the best-performing priors. We confirmed the applicability of our results to empirical data sets from cetaceans, phocids, and coregonid whitefish. Estimates of node times were generally robust to the choice of tree prior, but some combinations of tree priors and sampling schemes led to large differences in the age estimates. In particular, the pure-birth tree prior frequently led to inaccurate estimates for data sets containing a mixture of inter- and intraspecific sampling, whereas the birth-death and skyline coalescent priors produced stable results across all scenarios. Model testing provided an adequate means of rejecting inappropriate tree priors. Our results suggest that tree priors do not strongly affect Bayesian molecular dating results in most cases, even when severely misspecified. However, the choice of tree prior can be significant for the accuracy of dating results in the case of data sets with mixed inter- and intraspecies sampling.
Data from: How should genes and taxa be sampled for phylogenomic analyses with missing data? An empirical study in iguanian lizards
Targeted sequence capture is becoming a widespread tool for generating large phylogenomic data sets to address difficult phylogenetic problems. However, this methodology often generates data sets in which increasing the number of taxa and loci increases amounts of missing data. Thus, a fundamental (but still unresolved) question is whether sampling should be designed to maximize sampling of taxa or genes, or to minimize the inclusion of missing data cells. Here, we explore this question for an ancient, rapid radiation of lizards, the pleurodont iguanians. Pleurodonts include many well-known clades (e.g., anoles, basilisks, iguanas, and spiny lizards) but relationships among families have proven difficult to resolve strongly and consistently using traditional sequencing approaches. We generated up to 4921 ultraconserved elements with sampling strategies including 16, 29, and 44 taxa, from 1179 to approximately 2.4 million characters per matrix and approximately 30% to 60% total missing data. We then compared mean branch support for interfamilial relationships under these 15 different sampling strategies for both concatenated (maximum likelihood) and species tree (NJst) approaches (after showing that mean branch support appears to be related to accuracy). We found that both approaches had the highest support when including loci with up to 50% missing taxa (matrices with ∼40–55% missing data overall). Thus, our results show that simply excluding all missing data may be highly problematic as the primary guiding principle for the inclusion or exclusion of taxa and genes. The optimal strategy was somewhat different for each approach, a pattern that has not been shown previously. For concatenated analyses, branch support was maximized when including many taxa (44) but fewer characters (1.1 million). For species-tree analyses, branch support was maximized with minimal taxon sampling (16) but many loci (4789 of 4921). We also show that the choice of these sampling strategies can be critically important for phylogenomic analyses, since some strategies lead to demonstrably incorrect inferences (using the same method) that have strong statistical support. Our preferred estimate provides strong support for most interfamilial relationships in this important but phylogenetically challenging group.
Data from: Uneven sampling and the analysis of vocal performance constraints
Studies of trilled vocalizations provide a premiere illustration of how performance constraints shape the evolution of mating displays. In trill production, vocal tract mechanics impose a trade-off between syllable repetition rate and frequency bandwidth, with the trade-off most pronounced at higher values of both parameters. Available evidence suggests that trills that simultaneously maximize both traits are more threatening to males or more attractive to females, consistent with a history of sexual selection favoring high-performance trills. Here, we identify a sampling limitation that confounds the detection and description of performance trade-offs. We reassess 70 data sets (from 26 published studies) and show that sampling limitations afflict 63 of these to some degree. Traditional upper-bound regression, which does not control for sampling limitations, detects performance trade-offs in 33 data sets; yet when sampling limitations are controlled, performance trade-offs are detected in only 15. Sampling limitations therefore confound more than half of all performance trade-offs reported using the traditional method. An alternative method that circumvents this sampling limitation, which we explore here, is quantile regression. Our goal is not to question the presence of mechanical trade-offs on trill production but rather to reconsider how these trade-offs can be detected and characterized from acoustic data.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.