Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,694
datasets available to search
ShareScore release 0.7.1
Dataset results
4,694 results for “data analysis”
Raw data files and its analysis for MXene based PSCs.
<p>For more detailed information, one can visit <a title="Link to landing page via DOI" href="https://doi.org/10.1039/D4TC00466C" target="_blank" rel="noopener">10.1039/D4TC00466C</a></p> <p>Sanjay Sahare thanks project No. 2021/43/P/ST3/02599 co-funded by the National Science Centre and the European Union's Horizon 2020 research and innovation program under the Marie Skłodowska–Curie grant agreement no. 945339.</p>
Data and scripts for: "Exceptional point and hysteresis in perturbations of Kerr black holes" and "Massive scalar perturbations in Kerr Black Holes: near extremal analysis"
<p>Datasets associated with the article <strong>Exceptional point and hysteresis in perturbations of Kerr black holes </strong>- arXiv:2407.20850 [gr-qc] and <strong>Massive scalar perturbations in Kerr Black Holes: near extremal analysis</strong> - 2408.13964 [gr-qc]. The file Isomonodromy.zip contains three subfolders and the file CFM.zip contains two subfolders. Each subfolder includes a readme file that provides a description of the folder contents. A brief description of each subfolder is given below:</p> <p><strong>1) Isomonodromy.zip</strong></p> <p>1.1) Folder "Data" contains data obtained through the isomonodromic method. Data consists of radial eigenvalues (quasinormal frequencies) and angular eigenvalues as a function of the mass of the scalar field and the spin of the Kerr black hole. </p> <p>1.2) Folder "Notebook - Mathematica" contains a Mathematica notebook to compute the angular eigenvalue expansion and the asymptotic expression for the frequency in the extremal limit. </p> <p>1.3) Folder "Script - Julia" constains a script that implements the isomonodromic method to compute quasinormal modes of massive scalar perturbations in Kerr black holes.</p> <p>Note: This script has been tested with Julia 1.10.7 LTS and ArbNumerics 1.5.3. It is currently NOT compatible with Julia 1.11.X, likely due to changes in memory allocation breaking compatibility with libArb.</p> <p><strong>2) CFM.zip</strong></p> <p>1.1) Folder "Data" contains data obtained through the continued fraction method. Data consists of radial eigenvalues (quasinormal frequencies) and angular eigenvalues as a function of the mass of the scalar field and the spin of the Kerr black hole. </p> <p>1.2) Folder "Notebook - Mathematica" implements the continued fraction method to compute quasinormal modes of massive scalar perturbations in Kerr black holes.</p>
How to assess similarities and differences between mantle circulation models and Earth using disparate independent observations: Data and Analysis
<p>Dataset includes simulation output produced by a TERRA simulation for `How to assess similarities and differences between mantle circulation models and Earth using disparate independent observations'. </p> <p> </p> <h3><strong>Description of data file contents</strong></h3> <ul> <li><strong>NC*comp.tar.gz</strong> - compressed archives containing NetCDF files (file-per-process) with TERRA grid data including temperature, velocity, interpolated bulk composition, denisty, and voscosity fields. Can be read using <a title="terratools" href="https://github.com/mantle-convection-constrained/terratools" target="_blank" rel="noopener">terratools</a>. *dump number</li> <li><strong>NC_seis_037.tar.gz</strong> - compressed archive containing NetCDF files (file-per-process) with predicted seismic properties at the resolution of the TERRA grid generated from the present day state of the simulated mantle, including elastic and anelastic Vs and Vp, bulk sound velocity, and predicted density from mineral phyiscs tables. Can be read using <a title="terratools" href="https://github.com/mantle-convection-constrained/terratools" target="_blank" rel="noopener">terratools</a>.</li> <li><strong>NC_hpes_037.tar.gz</strong> - compressed archive containing NetCDF files (file-per-process) with interpolated abundances at the resolution of the TERRA grid for isotopes including the heat-producing elements ^40^K, ^232^Th, ^235^U and ^238^U. Can be read using <a title="terratools" href="https://github.com/mantle-convection-constrained/terratools" target="_blank" rel="noopener">terratools</a>.</li> <li><strong>P_files_037.tar.gz</strong> - compressed archive of TERRA P-files (particle files).</li> <li><strong>C_files_037.tar.gz</strong> - compressed archive of TERRA C-files (grid state files) - together with the P-files describe the full present day state of the simulation. </li> <li><strong>seis_filtered_037.tar.gz</strong> - compressed archive (file-per-layer) with reparameterised and seismically filtered (against S40RTS) present day Vs field.</li> <li><strong>seis_tables.tar.gz</strong> - compressed archive contianing lookup tables of seismic properties for the 3 principal lithologies assumed in the TERRA simulation (harzburgite, lherzolite and basaltic crust).</li> <li><strong>density_037.sph</strong> - Spherical harmonic coefficients for the density field in format to be read by the <a title="propagator" href="https://zenodo.org/records/12696774" target="_blank" rel="noopener">propagator matrix code.</a></li> <li><strong>plumes.pkl, ridges.pkl</strong> - Files containing tracer particle information for particles associated with plumes and ridges. </li> <li><strong>plumes_ridges.py</strong> - Python script containing example code for reading and plotting plumes.pkl and ridges.pkl files. </li> <li><strong>ptcls_rdgs_plms.py, interrogate_particles.py</strong> - Python script and module containing required functions for carrying out post processing routine generating the plumes.pkl and ridges.pkl files. Requires <a title="terratools" href="https://github.com/mantle-convection-constrained/terratools" target="_blank" rel="noopener">terratools</a>.</li> <li><strong>hst.dat </strong>- Time series of key simulation properties including mantle temperature profile used for calcualting CMB heat flux. </li> <li><strong>terra, interra</strong> - TERRA executable and input parameter file.</li> <li><strong>pyflowng.zip</strong> - Compressed directory containing version of the `pyflowng` code used in this work.</li> <li><strong>mode_splitting_methods.zip</strong> - Compressed directory contianing synthetic splitting function predictions and maps.</li> </ul> <p> </p> <h3><strong>Dump Numbers</strong></h3> <p>Below is a table of dump numbers (final three digits of file names) and the corresponding model times.</p> <table> <tbody> <tr> <td><strong>Dump number </strong></td> <td><strong>Model time (Ma)</strong></td> </tr> <tr> <td>037</td> <td>0 (present day)</td> </tr> <tr> <td>027</td> <td>10</td> </tr> <tr> <td>026</td> <td>20</td> </tr> <tr> <td>025</td> <td>30</td> </tr> <tr> <td>024</td> <td>40</td> </tr> <tr> <td>023</td> <td>50</td> </tr> <tr> <td>022</td> <td>60</td> </tr> <tr> <td>021</td> <td>70 </td> </tr> <tr> <td>020</td> <td>80</td> </tr> <tr> <td>019</td> <td>90</td> </tr> <tr> <td>018</td> <td>100</td> </tr> </tbody> </table>
Data for "Effects of forest dieback on deadwood patterns: large scale trends from a cross-analysis of European databases"
<p><strong><span>Aims</span></strong></p> <p><span>We carried out an opportunistic correlative study between past crown conditions and current deadwood volumes.</span></p> <p><span>Our aim was to mobilise available data on site factors and long-term monitoring of crown vitality indicators in Europe to investigate the influence of current and recent local defoliation levels on plot-level deadwood volume.</span></p> <p><span>For a subset of level I, 16*16-km monitoring plots located throughout Europe, we benefitted from data on both (i) deadwood measurements carried out within the framework of the Forest Focus Biosoil Project </span><span>(Galluzzi et al., 2019)</span><span>, pre-processed into a consistent and harmonized deadwood dataset by </span><span>Puletti et al. (2019)</span><span>, and (ii) defoliation assessments provided yearly since 1989 by the International Co-operative Program on Assessment and Monitoring of Air Pollution Effects on Forests (ICP Forests), the most comprehensive European monitoring network for the large-scale assessment of forest ecosystem health </span><span>(Vitale et al., 2014)</span><span>. </span></p> <p><span>Biosoil data on deadwood and ICP data on defoliation have never been crossed before.</span></p> <p><span>We used defoliation level as a proxy for the severity of stand dieback. Deadwood patterns can be addressed through deadwood profiles, which subdivide local deadwood stocks into classes based on size, position and decay stage.</span></p> <p><a name="_Toc175840512"></a><a name="_Toc116027761"></a><span><strong><span>ICP database and defoliation protocol</span></strong></span></p> <p><span>The International Cooperative Program to assess and monitor air pollution effects on the forest (ICP Forests) is responsible for an extensive level I monitoring system of forest sites </span><span>(Hauβmann & Fischer, 2004)</span><span>, which has been in operation since 1986. This large-scale level I network is made up of dense, spatially representative sampling points placed throughout European forests on a 16 × 16 km virtual grid, and is dedicated to monitoring forest conditions. The sampling points cover most European forested areas and encompasses ca. 6000 monitoring plots in 42 countries. In each plot, a visual evaluation of defoliation and discoloration of tree crowns is performed annually to survey forest health status (<a href="http://icp-forests.net/page/largescale-forest-condition">http://icp-forests.net/page/largescale-forest-condition</a>). Data management is presently carried out at the Programme Co-ordinating Centre (PCC) of ICP Forests in Eberswalde, Germany, and all data are available upon request. Since 1989, a standardized procedure for “annual surveys of crown condition’’ has been applied to 24 selected dominant and co-dominant trees with a minimum height of 60 cm and showing no significant mechanical damage. The defoliation and discoloration level of each tree crown is visually assessed on a sliding scale of 5% increments as the percentage of needle/leaf loss in the assessable crown as compared to a reference tree with full foliage. Mean defoliation at the plot scale was defined as the proportion of “damaged” trees i.e., with a defoliation rate of more than 25%, and used as a proxy for plot decline level. In the ICP database, the factors associated with observed defoliation related to natural disturbances or management (i.e., vertebrate or insect herbivory, fungal or fire damage, drought impacts, signs of removal of coarse woody debris, past landscape) were not recorded in a sufficiently standardized way to be used as covariates in our models. Similarly, plot-level living tree density and above-ground biomass for standing living trees (expressed in kg.ha<sup>−1</sup>), presumably surveyed in subplot 2, were not available.</span></p> <p><a name="_Toc175840513"></a><a name="_Toc116027762"></a><span><strong><span>Biosoil database and deadwood protocol</span></strong></span></p> <p><a name="_Toc116027763"></a><span>In the framework of the large collaborative European Forest Focus BioSoil-Biodiversity project</span><span>, a system of circular concentric subplots was built around certain ICP level I plots to collect additional data on stand structure and biodiversity between 2005 and 2008 (Figure 1). </span><span><span>The individual countries were responsible for selecting the ICP level I plots to be included in the BioSoil project </span></span><span><span>(Galluzzi et al., 2019)</span></span><span><span>. Overall, a total of 3243 geocoded Level I plots were considered in 19 European countries </span></span><span><span>(Puletti et al., 2017)</span></span><span><span>: Austria, Belgium (Flanders only), Cyprus, the Czech Republic, Denmark, Finland, France, Germany (eight federal states only), Hungary, Ireland, Italy, Latvia, Lithuania, Poland, Slovakia, Slovenia, Spain, Sweden and the United Kingdom (Figure 1). BioSoil project results are recorded in the multi-dimensional LI-BioDiv geodatabase that contains raw data on forest structure and vegetation records used to calculate simple plot-level structural and compositional forest variables (i.e., biomass, deadwood volume, plant alpha-diversity; </span></span><span><span>Bastrup-Birk et al. 2007; Hiederer & Durant 2010)</span></span><span><span>. At each plot, deadwood was quantified on an area of 400 m<sup>2</sup> (BioSoil subplots 1 and 2, radius of 11.28 m; </span></span><span><span>Puletti et al., 2017)</span></span><span><span>. The deadwood survey included coarse woody debris (including lying dead trees), snags (including standing dead trees) and stumps more than 10 cm in diameter. Only snags and stumps more than 130 cm in height were considered. Diameter, length or height, tree species and decay stage (5 classes) were recorded for each deadwood piece. The raw ICP deadwood data were processed by </span></span><span><span>Puletti et al. (2017, 2019)</span></span><span><span> into a consistent and harmonized pan-European deadwood dataset, which we used in this study. The dataset provides total deadwood volume and the volume of several deadwood types for each plot. Further details can be found in the ICP Forests manual (</span></span><a href="http://icp-forests.net/page/icp-forests-manual"><span><span>http://icp-forests.net/page/icp-forests-manual</span></span></a><span><span>), </span></span><span><span>Puletti et al. (2019)</span></span><span><span> and </span></span><span><span>Augustynczik et al. (2024)</span></span><span><span>.</span></span></p> <p><span><span>In our study, we considered the following response variables</span></span><span>: (i) total deadwood volume, (ii) </span><span>standing deadwood (snags) volume, (iii) volume of ground-lying deadwood, (iv) </span><span>fresh deadwood volume </span><span>(= Vm3_dec1_Biosoil + Vm3_dec2_Biosoil), and (v) decayed deadwood volume = (= Vm3_dec4_Biosoil + Vm3_dec5_Biosoil).</span></p> <p><span>A few environmental covariates were collected from the Biosoil data: (i) management intensity (grouped into two classes: recently harvested, i.e., with management evidence within the last 10 years; and not recently harvested, i.e., unmanaged (no management evidence) or managed a long time ago (management evidence but more than 10 years previously), (ii) average stand age (separated into 3 classes: mature [>100 yrs], mid-aged [41-100 yrs], young [1-40 yrs]), (iii) elevation (above sea level, a.s.l.), a continuous quantitative variable, (iv) dominant tree genus, and (v) forest type, depending on the dominant tree species: coniferous, deciduous or mixed.</span></p> <p><a name="_Toc175840514"></a><a name="_Toc116027764"></a><span><strong><span>Database joint</span></strong></span><span><strong><span>: <a name="_Toc116027765"></a>plot matching in time series</span></strong></span></p> <p><span>After harmonizing plot names and coordinates in the two datasets (ICP-defoliation and Biosoil-deadwood), only plots with matched data in both datasets were selected. Plots with a maximum of one year’s discontinuity in the data were retained, and the missing values were reconstructed from the average values in contiguous years. Plots with discontinuities in defoliation measurements of more than 2 years were deleted. We matched defoliation measurements for the Biosoil-ICP datasets from 1989 to 2007 and finally obtained 2,070 five-year, 1,804 ten-year and 1,399 fifteen-year time series. This approach made it possible to define three 10-year time series [1995-2005, 1996-2006, 1997-2007] with plots in 17 countries, from five plots in Ireland and nine in the United Kingdom, to 337 plots in Finland and 461 in France.</span></p> <p><a name="_Toc175840515"></a><a name="_Toc116027766"></a><span><strong><span>Calculation of global defoliation metrics</span></strong></span></p> <p><span>We calculated 16 univariate metrics to summarize changes in defoliation throughout the 10-year period prior to the Biosoil deadwood measurements. Some of the selected parameters describe the immediate possible effects of defoliation severity in the recent past on a given year: (i) defoliation level of the previous year (n-1), (ii) defoliation level of the year before the previous year (n-2), (iii) defoliation level of the year two years before the previous year (n-3). Other defoliation metrics relate to the cumulative effects of defoliation levels in the near or the distant past: (i) average defoliation level over the last two years, (ii) average defoliation level over the last three years, (iii) average defoliation level over the last five years, (iv) average defoliation level over the first five years of the 10-year time series, and (v) time elapsed since last peak defoliation. Several other parameters depict general trends in the level of defoliation over the 10-year time series: for cumulative metrics: (i) arithmetic mean of annual defoliation level; (ii) geometric mean of annual defoliation level; (iii) Area Under the defoliation time Curve (AUC), i.e., the cumulative sum of defoliation levels; and for the overall trend: (iv) the estimated slope of the linear regression line for defoliation level over time. Finally, some of the metrics reflect defoliation severity and repetition along the 10-year time series, and their potentially time-lagged effects: (i) maximum defoliation level; (ii) total number of years elapsed after the dieback peak level, whether successive or not; (iii) the number of peaks, consecutive or discontinuous, i.e., the number of severe defoliation events and defoliation frequency; and (iv) duration of the longest peak, i.e., the longest continuous time during which the level of defoliation was greater than the relative threshold.</span></p> <p><span>A peak in defoliation was defined as a year in which the level of defoliation exceeded a relative threshold, i.e., the third quartile value. In our 10-year time series, the peak value was 25% and above. <span><span> </span></span></span></p>
Data and Analysis Files Repository: Repurposing Large-Format Microarrays for Scalable Spatial Transcriptomics
<p>Data and Analysis Files from "Repurposing Large-Format Microarrays for Scalable Spatial Transcriptomics"</p> <p>ArraySeq_Method.zip contains the following folder and contents:</p> <ul> <li>STARSolo: All code and count matrix output from fastq spatial barcode demultiplexing. </li> <li>Images: All resolution-downsampled H&E image scans from analyzed tissues</li> <li>Space_Ranger: All 10x Space Ranger output from Visium datasets generated in the paper. </li> <li>Analysis: All scripts for analyzing and plotting Array-seq and Visium datasets generated in this paper. Also contains output h5ad files. </li> </ul> <p>ArraySeq_Barcode_generation_n12.rmd: The script used to generate the Array-seq probes with 12-mer spatial barcodes. </p>
Neutron activation analysis data of pottery from El Rincon, Los Angeles, Martinez, Piedras Blancas, Bordo de los Indios, and other sites, Argentina
<p>Neutron activation analysis data of pottery from El Rincon, Los Angeles, Martinez, Piedras Blancas, Bordo de los Indios, and other sites in Catamarca Province, Argentina</p>
Neutron activation analysis data of pottery from Rio Grande, Parque Grand Chaco, and Bañados del Izozog, Bolivia
<p>Neutron activation analysis data of pottery from Rio Grande, Parque Grand Chaco, and Bañados del Izozog, Santa Cruz Department, Bolivia.</p>
Intermediate Data Files for Amygdala and Posterior Piriform Cortex RNA-ATAC Analysis
<p>### Description of Data Files</p> <p>This dataset contains intermediate data files generated for the analysis of RNA and ATAC sequencing data from the amygdala and posterior piriform cortex. The files include ATAC-seq peak data, LOESS regression results, differential expression analysis, Gene Ontology (GO) analysis results, marker gene lists, and Seurat objects for further integrative analysis. Scripts and computational workflow can be found at https://github.com/karmaout/SOC_Multiomic_sequencing.</p> <table> <tbody> <tr> <th><strong>Filename</strong></th> <th><strong>Description</strong></th> </tr> <tr> <td>PC_PR_Pyramidal_cfos_positive_atac.bed</td> <td>BED file containing ATAC-seq peaks for cfos-positive VGLUT1a neurons in the PC under PR conditions.</td> </tr> <tr> <td>BLA_PR_VGLUT1_cfos_positive_atac.bed</td> <td>BED file with ATAC-seq peaks for cfos-positive VGLUT1 neurons in the BLA under PR conditions.</td> </tr> <tr> <td>PC_PE_Pyramidal_cfos_negative_atac.bed</td> <td>BED file for ATAC-seq peaks of cfos-negative VGLUT1a neurons in the PC under PE conditions.</td> </tr> <tr> <td>PC_PR_loess_regression_data.csv</td> <td>CSV file containing data for LOESS regression analysis in the PC under PR conditions.</td> </tr> <tr> <td>BLA_cfos_percentages_raw_data.csv</td> <td>Raw data of cfos expression percentages in different neuron populations in the BLA.</td> </tr> <tr> <td>PC_PE_loess_regression_data.csv</td> <td>LOESS regression data for the PC under PE conditions.</td> </tr> <tr> <td>PC_neuron_cluster_markers.csv</td> <td>CSV file listing marker genes for PC neuron clusters.</td> </tr> <tr> <td>BLA_PE_VGLUT1_cfos_positive_atac.bed</td> <td>BED file for ATAC-seq peaks of cfos-positive VGLUT1 neurons in the BLA under PE conditions.</td> </tr> <tr> <td>BLA_PR_VGLUT1_cfos_negative_atac.bed</td> <td>BED file for ATAC-seq peaks of cfos-negative VGLUT1 neurons in the BLA under PR conditions.</td> </tr> <tr> <td>BLA_Subcluster_Differential_Expression_Results(with Fos).xlsx</td> <td>Differential expression results for BLA subclusters with Fos gene.</td> </tr> <tr> <td>PC_Aggregation.csv</td> <td>Aggregated data for PC neurons, summarizing key statistics or metrics.</td> </tr> <tr> <td>PC_Subcluster_Differential_Expression_Results_PC(with Fos).xlsx</td> <td>Differential expression results for PC subclusters with Fos gene.</td> </tr> <tr> <td>PC_cfos_percentages_raw_data.csv</td> <td>Raw percentages of cfos expression in different PC neuron subclusters.</td> </tr> <tr> <td>PC_PE_Pyramidal_cfos_positive_atac.bed</td> <td>BED file for ATAC-seq peaks of cfos-positive Pyramidal neurons in the PC under PE conditions.</td> </tr> <tr> <td>PC_PR_Pyramidal_cfos_negative_atac.bed</td> <td>BED file for ATAC-seq peaks of cfos-negative Pyramidal neurons in the PC under PR conditions.</td> </tr> <tr> <td>BLA_PR_loess_regression_data.csv</td> <td>LOESS regression analysis data for the BLA under PR conditions.</td> </tr> <tr> <td>BLA_GO_analysis_results_all_combined.xlsx</td> <td>Combined Gene Ontology (GO) analysis results for BLA neuron populations.</td> </tr> <tr> <td>PC_GO_analysis_results_all_combined.xlsx</td> <td>Combined GO analysis results for PC neuron populations.</td> </tr> <tr> <td>BLA_PR_anno_sorted_foldchange.txt</td> <td>Annotated fold change data for BLA under PR conditions.</td> </tr> <tr> <td>BLA_PE_anno_sorted_foldchange.txt</td> <td>Annotated fold change data for BLA under PE conditions.</td> </tr> <tr> <td>PC_PR_anno_sorted_foldchange.txt</td> <td>Annotated fold change data for PC under PR conditions.</td> </tr> <tr> <td>PC_PE_anno_sorted_foldchange.txt</td> <td>Annotated fold change data for PC under PE conditions.</td> </tr> <tr> <td>BLA_PE_loess_regression_data.csv</td> <td>Data from LOESS regression analysis in the BLA under PE conditions.</td> </tr> <tr> <td>BLA_PE_VGLUT1_cfos_negative_atac.bed</td> <td>BED file for ATAC-seq peaks of cfos-negative VGLUT1 neurons in the BLA under PE conditions.</td> </tr> <tr> <td>BLA_neuron_cluster_markers.csv</td> <td>Marker genes for various neuron clusters in the BLA.</td> </tr> <tr> <td>BLA_Aggregation.csv</td> <td>Aggregated data for BLA neurons, summarizing key metrics or statistics.</td> </tr> <tr> <td>BLA_neurons.rds</td> <td>RDS file of the Seurat object for BLA neurons containing RNA and ATAC analysis.</td> </tr> <tr> <td>PC_neurons.rds</td> <td>RDS file of the Seurat object for PC neurons with integrated RNA and ATAC data.</td> </tr> <tr> <td>BLA_matrix-scan.ft.txt</td> <td>Matrix scan results for motif analysis in the BLA.</td> </tr> <tr> <td>PC_coord_anno.txt</td> <td>Annotated genomic coordinates for PC, used for integrative analysis.</td> </tr> <tr> <td>BLA_coord_anno.txt</td> <td>Annotated genomic coordinates for BLA.</td> </tr> <tr> <td>PC_matrix-scan.ft.txt</td> <td>Matrix scan results for motif analysis in the PC.</td> </tr> </tbody> </table>
Monte Carlo Analysis Code and Raw Data
Open the record for dataset details and reuse information.
Complex ecological phenotypes on phylogenetic trees: a Markov process model for comparative analysis of multivariate count data
The evolutionary dynamics of complex ecological traits – including multistate representations of diet, habitat, and behavior – remain poorly understood. Reconstructing the tempo, mode, and historical sequence of transitions involving such traits poses many challenges for comparative biologists, owing to their multidimensional nature. Continuous-time Markov chains (CTMC) are commonly used to model ecological niche evolution on phylogenetic trees but are limited by the assumption that taxa are monomorphic and that states are univariate categorical variables. A necessary first step in the analysis of many complex traits is therefore to categorize species into a pre-determined number of univariate ecological states, but this procedure can lead to distortion and loss of information. This approach also confounds interpretation of state assignments with effects of sampling variation because it does not directly incorporate empirical observations for individual species into the statistical inference model. In this study, we develop a Dirichlet-multinomial framework to model resource use evolution on phylogenetic trees. Our approach is expressly designed to model ecological traits that are multidimensional and to account for uncertainty in state assignments of terminal taxa arising from effects of sampling variation. The method uses multivariate count data for individual species to simultaneously infer the number of ecological states, the proportional utilization of different resources by different states, and the phylogenetic distribution of ecological states among living species and their ancestors. The method is general and may be applied to any data expressible as a set of observational counts from different categories.
Data from: Forecasting potential emergence of zoonotic diseases in Southeast Asia: network analysis identifies key rodent hosts
1. Within complex ecological systems, identifying animal species likely to play a key role in the emergence of infectious zoonotic diseases remains a major challenge. One approach consists of using information on current ecological and parasitological similarities among host species in order to predict the most likely pathways for future pathogen spillover. 2. Using field data acquired from 15 sympatric rodent species in various habitats in Thailand, Cambodia and Laos, we built networks based on shared parasites (17 helminth and 15 microparasite species) and shared habitats among rodent species and humans. We investigated the architectures of bipartite and unipartite networks using modularity, subgroups partitioning or node centrality, to assess the relative epidemiological importance of particular rodent species. 3. Our results showed that Rattus tanezumi, Bandicota savilei and R. exulans were consistently found to be members of subgroups that included humans in unipartite and bipartite networks on zoonotic agents and shared habitats. High values of centrality in shared zoonotic agents were found for the same three rodent species, whereas high values of shared habitats were observed for two of them. Although phylogenetically related rodent species likely shared both habitats and parasites, a lack of habitat specialisation was associated with increased zoonotic parasite sharing. 4. Our results emphasize the disproportionate importance of these three rodent species, through their high degree of connectivity with humans, which may represent a high risk for direct zoonotic spillover. Moreover, due to its high centrality in habitats, R. tanezumi may also play a key role as a bridge host. 5. The recent discovery of new arenaviruses in rodents in Southeast Asia, with associated disease in humans in Cambodia, provides an opportunity to test this empirically. The three rodent species identified using our network approach are some of the potential maintenance hosts for these new emerging arenaviruses. 6. Synthesis and applications. Our results on rodents and their pathogens in Southeast Asia show that network analysis has a high potential to improve the surveillance of emerging zoonotic pathogens by targeting key host species and potential "emerging' pathogen–rodent interactions in complex and heterogeneous landscapes.
Data Matrix Theme-Specific Analysis of the Recommendation on Science and Scientific Researchers (RSSR): Open Access, Open Data, and Open Science
<p>This Table sets out findings from the mapping exercise conducted as part of the objectives of subtask 6.1 of the RRING project.</p> <p>Aim: Alignment of RRI to advance the UN SDGs.</p> <p>Objectives:</p> <ul> <li>Mapping the RSSR to the SDGs </li> </ul> <p>Mapping the RSSR to the SDGs is aimed at providing new perspectives, ideas and approaches that can help to improve the operationalization and implementation of each SDG, <em>by facilitating the integration of RRI (or RRI-like) practices in the SDGs, to make them more achievable.</em> The impact of the new perspectives, ideas and approaches in SDG operationalization and implementation will be aimed at the level of <em>national and international policy (making); future research and innovation projects (in industry and academia); as well as education and training of researchers, policy makers and other stakeholders.</em></p> <p>Two documents were used for this task:</p> <ul> <li>2017 Recommendation on Science and Scientific Researchers ([RSSR], UNESCO), and</li> <li>the United Nations 2030 Agenda for Sustainable Development with the 17 Sustainable Development Goals (SDGs).</li> </ul>
Data and analysis scripts for "Mad7: An IP friendly alternative to Cas9"
<p>This is the raw data and analysis scripts for "Mad7: An IP friendly alternative to Cas9"</p>
Data from: Whole-genome analysis of Mustela erminea finds that pulsed hybridization impacts evolution at high-latitudes
At high-latitude, climatic shifts hypothetically drove episodes of divergence during isolation in glacial refugia, or ice-free pockets of land that enabled terrestrial species persistence. Upon glacial recession, populations can expand and often come into contact, resulting in admixture between previously isolated groups. To understand how recurrent periods of isolation and contact have impacted evolution at high latitudes, we investigated introgression in the stoat (Mustela erminea), a Holarctic mammalian carnivore, using whole-genome sequences. We identify two temporally isolated introgression events coincident with large-scale climatic shifts: contemporary introgression in a mainland contact zone and ancient contact ~ 200 km south along North America's North Pacific Coast. Repeated episodes of gene flow highlight the central role of cyclic climates in structuring high-latitude diversity, through refugial divergence and subsequent introgressive hybridization. Introgression followed by allopatry (e.g., insularization) may contribute to expedited divergence of island taxa experiencing substantial glacial flux.
Data sets of the analysis of 1sg and 2sg subject expression with the verbs 'creer' and 'saber' in a corpus of spoken Spanish
<p>These are the two data sets used for the quantitative analysis in the paper "“Perspectival factors and <em>pro</em>-drop – A corpus study of speaker/addressee pronouns with <em>creer </em>‘think/believe’ and <em>saber </em>‘know’ in spoken Spanish” (Peter Herbeck;<em> to appear </em>in <em>Glossa – a journal of general linguistics</em>). It contains the values of the annotation of null and overt speaker/addressee pronouns in the Madrid and Alcalá samples of the corpus PRESEEA (2014-). The first data set was used for the quantitative analysis of subject expression (null/overt) with the verbs <em>creer </em>and <em>saber </em>according to person (1sg vs. 2sg) and polarity (negative vs. positive verb forms). The second data set was used for the analysis of 1sg subject expression according to the complement type of <em>creer </em>and <em>saber</em>.</p>
Images, data, and statistical analysis scripts for review article on cover crop roots
<p>Images, data, and statistical analysis scripts for review article on cover crop roots.</p> <blockquote> <p><strong>Optimization of root traits to provide enhanced ecosystem services in agricultural systems: a focus on cover crops</strong> - [<a href="https://doi.org/10.1111/pce.14247">https://doi.org/10.1111/pce.14247</a>]</p> </blockquote> <ul> <li>Research site, planting, and growth <ul> <li>10/2020 - 04/26/2021 cover crop field trial. DDPSC FRS at Planthaven Farm, O'Fallon, MO 63366 (latitude 38.848240°, longitude -90.686640°). </li> <li>The field was tilled before sowing of cover crops. Seed for each cover crop were spread in using a push seed spreader and were lightly irrigated.</li> <li>Alfalfa (<em>Medicago sativa</em>), dundale pea (<em>Pisum sativum</em>), milkvetch (<em>Astragalus canadensis</em>, <em>Astragalus bisulcatus</em>), crimson clover (<em>Trifolium incarnatum</em>), hairy vetch (<em>Vicia villosa</em>), mustard (<em>Brassica junce</em>a var Mighty Mustard, var Kodiak), barley (<em>Hordeum vulgare</em>), wheat (<em>Triticum aestivum</em>, winter, spring), winter rye (<em>Secale cereale</em>), and triticale (× T<em>riticosecale</em> Wittmack).</li> </ul> </li> <li> <p>Field harvest measurements</p> <ul> <li> <p>Four canopy images were taken across each cover crop row using a Canon 5DS R camera. Images were taken from above each plot at 5ft height manually. Green color was thresholded from the canopy images in batch using OpenCV python script and the percent green cover calculated (Jupiter notebook).</p> </li> <li> <p>Five soil monoliths were excavated using a "shovelomics" approach with an average monolith size of 25.4cm x 25.4cm x 20 cm. The remaining four soil monoliths were destructively analyzed.</p> </li> <li> <p>One soil monolith was imaged using a Canon 50D DLSR camera in a photogrammetry shed. All photogrammetric analysis was conducted using Pix4D mapper software (Pix4D S.A. Prilly, Switzerland), and point cloud cleaning was conducted in CloudCompare V2. 10.2.</p> </li> <li> <p>Cover crop shoots from the remaining soil monoliths were cut and placed into a paper bag for dry biomass determination (60oC for 5 days). A cover crop shoot count was conducted for each monolith with each tiller considered as a shoot for the grasses (barley, wheat, triticale). After cover crop shoot harvesting, a photo was then taken of each soil monolith with remaining weed biomass. A weed score was assigned to each image by one trained researcher with a score 1 low weeds to 5 high weed presence.</p> </li> <li> <p>Soil monoliths were the soaked briefly in water and then the soil washed using a hose keeping the roots. Roots were then scanned on an Epson Expression 12000XL Photo Scanner with transparency unit. Images labeled with "_part" were samples with too many roots for scanning and so were separately weighed. Dry root biomass was taken for the scanned and unscanned roots separately. Root length was determined from images using software RhizoVision Explorer (https://doi.org/10.5281/zenodo.4095629), total root length was estimated using scanned root length and scanned dry biomass with unscanned root biomass.</p> </li> <li> <p>Along each cover crop plot a 10ft trench was dug using a Yanmar Excavator Vi020-6 perpendicular to the row with each trench fully bisecting the plot. Trench was one bucket wide (19 inches) and approximately 36 inches deep in the middle of the row. The five deepest roots that could be observed in the trench wall was measured manually with a tape measure for each cover crop. A garden trowel and shovel were used to excavate and confirm roots in trench wall.</p> </li> <li> <p>Data was analyzed using R Statistics script and raw data used for data processing and figure generation (2021PlantHavenCovercrop_dataprocessing.R). PCA analysis was conducted using the “FactoMineR” package (Husson <em>et al</em>. 2019) to explore the relationships between the traits within the dataset and clustered by family.</p> </li> </ul> </li> </ul> <p>Individual ZIP file contents:</p> <ul> <li><code><strong>2021PlantHavenCovercrop_CanopyImages.zip</strong></code> – Raw canopy images, processed percent green cover images, and Jupiter notebook python script (2021PlantHavenCovercrop_ImageBatchColorThreshold.ipynb).</li> <li><code><strong>2021PlantHavenCovercrop_RootFlatbedImages.zip</strong></code> – Raw flatbed root scans of cover crops and processed images using RhizoVision Explorer.</li> <li><code><strong>2021PlantHavenCovercrop_SoilMonolithWeedImages.zip</strong></code> – Images of soil monoliths after cover crop shoot biomass was removed.</li> <li><code><strong>2021PlantHavenCovercrop_dataprocessing.zip</strong></code> – R Statistics script and raw data used for data processing and figure generation (2021PlantHavenCovercrop_dataprocessing.R).</li> <li><code><strong>2021PlantHavenCovercrop_ShootPhotogrammetry.zip</strong></code> – 3D models of cover crop shoots from excavated soil monoliths. The .bin files can be opened using CloudCompare app.</li> </ul> <p> </p> <p> </p>
Data from: Investigating cat predation as the cause of bat wing tears using forensic DNA analysis
<p>Cat predation upon bat<i> </i>species has been reported to have significant effects on bat populations in both rural and urban areas. The majority of research in this area has focussed on observational data from bat rehabilitators documenting injuries, and cat owners, when domestic cats present prey. However, this has the potential to underestimate the number of bats killed or injured by cats. Here, we use forensic DNA analysis techniques to analyse swabs taken from injured bats in the United Kingdom, mainly including <i>Pipistrellus pipistrellus </i>(40 out of 72 specimens)<i>. </i>Using quantitative PCR, cat DNA was found in two-thirds of samples submitted by bat rehabilitators. Of these samples, short tandem repeat analysis produced partial DNA profiles for approximately one-third of samples, which could be used to link predation events to individual cats. The use of genetic analysis can complement observational data, and potentially provide additional information to give a more accurate estimation of cat predation. </p>
Data from: Comparative genomic analysis of the pheromone receptor Class 1 family (V1R) reveals extreme complexity in mouse lemurs (genus, Microcebus) and a chromosomal hotspot across mammals
<p><span>Sensory gene families are of special interest, both for what they can tell us about molecular evolution, and for what they imply as mediators of social communication. The vomeronasal type-1 receptors (V1Rs) have often been hypothesized as playing a fundamental role in driving or maintaining species boundaries given their likely function as mediators of intraspecific mate choice, particularly in nocturnal mammals. Here, we employ a comparative genomic approach for revealing patterns of V1R evolution within primates, with a special focus on the small-bodied nocturnal mouse and dwarf lemurs of Madagascar (genera <i>Microcebus</i> and <i>Cheirogaleus</i>, respectively). By doubling the existing genomic resources for strepsirrhine primates (i.e., the lemurs and lorises), we find that the highly speciose and morphologically cryptic mouse lemurs have experienced an elaborate proliferation of V1Rs that we argue is functionally related to their capacity for rapid lineage diversification. Contrary to a previous study that found equivalent degrees of V1R diversity in diurnal and nocturnal lemurs, our study finds a strong correlation between nocturnality and V1R elaboration, with nocturnal lemurs showing elaborate V1R repertoires and diurnal lemurs showing less diverse repertoires. Recognized subfamilies among V1Rs show unique signatures of diversifying positive selection, </span>as might be expected if they have each evolved to respond to specific stimuli<span>. Further, a detailed syntenic comparison of mouse lemurs with mouse (genus <i>Mus</i>) and other mammalian outgroups shows that orthologous mammalian subfamilies, predicted to be of ancient origin, tend to cluster in a densely populated region across syntenic chromosomes that we refer to as a V1R "hotspot."</span></p>
Data Appendix for Lack, P., "Using Word Analysis to Track the Evolution of Emotional Well-being in Nineteenth-Century Industrializing Britain", Historical Methods (forthcoming)
<p>This file contains the data associated with the publication Lack, P., "Using Word Analysis to Track the Evolution of Emotional Well-being in Nineteenth-Century Industrializing Britain", <em>Historical Methods</em> (forthcoming). It quantifies the trend in emotional well-being expressed in a corpus of British pamphlets published between 1800 and 1900. The first page of the excel document presents this key data on the trend in emotional well-being. Sheet 1A presents summary statistics on the trend in emotional well-being and its correlation with GDP per capita and real wages. </p>
Simulation and analysis data set for apo-conformational kinetics and gated ligand binding to HIV-1 protease
<p>The data set provided here accompanies a study described in the manuscript:</p> <p>S. Kashif Sadiq, Abraham Muñiz Chicharro, Patrick Friedrich, Rebecca Wade, A multiscale approach for computing gated ligand binding from molecular dynamics and Brownian dynamics simulations. (2021) Preprint available: https://doi.org/10.1101/2021.06.22.449380</p> <p>This study combines molecular dynamics MD simulations and associated conformational analyses and Markov state models (MSMs) with Brownian dynamics (BD) simulations to compute conformation gated ligand association kinetics to HIV-1 protease.</p> <p>To download the data, go to a directory where you would like to download the files. Then for each of the provided tar files enter the following command:</p> <p>tar xvf $X.tar</p> <p>where $X is the name prefix of the corresponding tar file.</p> <p>The unpacked data set creates a ./data sub-directory which itself contains two further sub-directories: MD and BD. Please see README.txt files within these sub directories for further instructions on the software tools and scripts that have been provided therein for using and reproducing the data set. The MD README.txt is found within: data_MD_MSM_analysis.tar, the BD README.txt is found within: data_BD_examples.tar.</p> <p>Please note, the python Jupyter notebook and associated module for further analysis of the MSM from the pre-defined feature set calculated in the study as well as other analyses can also be found at:</p> <p>https://github.com/kashifsadiq/hiv1pr-msm/</p> <p>MD trajectory files are provided for further analysis but are not required to reproduce the MSM and conformational analyses reported in the study. To facilitate overview, MSM analysis has been stored in several object files. To exactly reproduce the reported MD/MSM analyses, untar only the 1) data_MD_MSM_analysis.tar and 2) data_MD_MSMobj.tar files and work through the python Jupyter notebook.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.