Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
409
datasets available to search
ShareScore release 0.9.0
Dataset results
409 results for “information use”
Simulation output for Improving the stability of bivariate correlations using informative Bayesian priors: A Monte Carlo simulation study
<p>This repository contains the (compressed) simulation output from <em>Improving the stability of bivariate correlations using informative Bayesian priors: A Monte Carlo simulation study</em>. On a Linux-based system, extract the contents with:</p> <pre><code class="language-bash">tar -xvzf raw_output_compressed.tar.gz </code></pre> <p>Please refer to the published article (https://doi.org/10.3389/fpsyg.2023.1253452) and the associated GitHub (<a href="https://github.com/carldelfin/stability-of-bivariate-correlations">github.com/carldelfin/stability-of-bivariate-correlations</a>) for additional information.</p>
TrainTicket microservice testbench extracted information for our work: Evaluating ChatGPT's Proficiency in Understanding and Answering Microservice Architecture Queries Using Source Code Insights
<p>It contains the CSV file output of our tool implemented in the paper: "Evaluating ChatGPT’s Proficiency in Understanding and Answering Microservice Architecture Queries Using Source Code Insights." applied to the TrainTicket microservice testbench. The information in this CSV was used for In-Context-Learning for ChatGPT.</p>
Data for The impact of information about tobacco-related reproductive vs. general health risks on South Indian women's tobacco use decisions
<p>Tobacco Intervention Study Mysore India March-April 2016</p> <p>Published version: <a href="https://doi.org/10.1017/ehs.2020.61">https://doi.org/10.1017/ehs.2020.61</a></p>
Full information on the eORCA1 grid (mesh_mask) used in IPSL-CM6A-LR configuration
<p>eORCA1.2_mesh_mask.nc : This file contains all relevant information on the eORCA1 grid used in the NEMO_v3.6_STABLE configuration of the oceanic module of the IPSL-CM6A-LR climate model. See https://www.nemo-ocean.eu/wp-content/uploads/NEMO_book.pdf for more details on the grid.</p> <p>eORCA_R1_bathy_meter_v2.2.nc: This file contains the bathymetry of the eORCA1 configuration used in IPSL-CM6A-LR.</p>
Supplementary Information S1 - Detailed results of the CAPRI N-LCA and S2 - Quantification of the main N budget flows in the EU25 agriculture sector of Leip, A., Billen, G., Garnier, J., Grizzetti, B., Lassaletta, L., Reis, S., Simpson, D., Sutton, M. a, de Vries, W., Weiss, F., Westhoek, H. (2015). Impacts of European livestock production: nitrogen, sulphur, phosphorus and greenhouse gas emissions, land-use, water eutrophication and biodiversity. Environ. Res. Lett. 10, 115004. doi:10.1088/1748-9326/10/11/115004
<p>Table S1-1 Quantification of GHG and Nr flow intensities [kg CO2eq (kg product)<sup>-1</sup> yr<sup>-1</sup>] or [g N (kg product)<sup>-1</sup> yr<sup>-1</sup>] with the CAPRI N-LCA model for six main livestock products (BEEF: beef, PORK: pork, EGGS: eggs, POUM: poultry meat; DAIR: milk and dairy products, SGMP: meat from sheep and goats) and six main vegetable food groups (POTA: potatoes, SUGB: sugar beet before processing, OILP: oil seeds before processing; CERR: cereals, LEGU: leguminous crops) as well as other crops (OCRP) and aggregated livestock (ANIMP) and vegetable (CROPP) food. </p> <p>Table S2-1 Quantification of the main N budget flows in the EU25 agriculture sector</p>
Can accurate demographic information about people who use prescription medications non-medically be derived from Twitter?
<p>This archive contains over 3 billion Tweet IDs associated with the paper:<br>"Can accurate demographic information about people who use prescription medications non-medically be derived from Twitter?"</p> <p>The data provides:<br>X (Twitter) IDs of posts included in the study. The IDs can be used to retrieve the original posts via the Twitter API. Posts removed by the original subscribers or whose visibilities are no longer public cannot be retrieved by the poster. </p> <p>Full citation:<br>Yang YC, Al-Garadi MA, Love JS, Cooper HLF, Perrone J, Sarker A. Can accurate demographic information about people who use prescription medications nonmedically be derived from Twitter? Proc Natl Acad Sci U S A. 2023 Feb 21;120(8):e2207391120. doi: 10.1073/pnas.2207391120. Epub 2023 Feb 14. PMID: 36787355; PMCID: PMC9974473.</p> <p>Python scripts related to the analysis are available as supplementary material with the paper. </p> <p>Contact: <br>Abeed Sarker<br>abeed@dbmi.emory.edu</p> <p>Funding:<br>National Institute on Drug Abuse (R01DA057599).</p>
Dispilio. Supplementary Information for Maczkowski et al., Absolute dating of the European Neolithic using the 5259 BC rapid 14C excursion
<p>Supplementary Material for the paper "Absolute dating of the European Neolithic using the 5259 BC rapid 14C excursion":</p> <p> </p> <p><strong>Supplementary Information</strong> includes OxCal code, wiggle-matching output, photographs of the Neolithic juniper wood samples analysed, photographs of the tree-ring sampling for annual radiocarbon, photographs of modern tree-ring analogues, supplementary text on the data presented in the article, as well as the tree-ring width measurements in Heidelberg format (.fh).</p> <p><strong>Supplementary Data 1-2 </strong>includes spreadsheets with all the new raw radiocarbon data presented in the article, the associated uncertainties and ring numbers.</p> <p><strong>Supplementary Data 3</strong> includes the R code and the source data used for the generation of Figures 3 and 5 in the main article text, as well as the OxCal code used for the wiggle-matching of annual 14C in OxCal as presented in Fiugre 5</p> <p>The latest version of the Supplementary Material is just an expanded version of the first, files have been renamed according to editorial guidlines, few extra figures, OxCal code, and extra information added after the review process. No changes were made to any of the data published online in the initial version of the Supplementary Material.</p> <p> </p> <p>File renaming from last version:</p> <p>Supplementary Material = Supplementary Information</p> <p>Supplementary Material S1 = Supplementary Note 1</p> <p>Supplementary Material S2 = Supplementary Figures</p> <p>Supplementary Material S3 = Supplementary Note 2</p> <p>Supplementary Table T1 = Supplementary Data 1-2</p> <p>Supplementary Material S4 = Supplementary Data 3</p> <p> </p>
Improved Bathymetric Prediction using Geological Information: SYNBATH
<p>Manuscript in revision: <em>Earth and Space Science, </em>December 20, 2021</p> <p><em>Abstract</em></p> <p>To date, approximately 20% of the ocean floor has been surveyed by ships at a spatial resolution of 400 m or better. The remaining 80% has depth predicted from satellite altimeter-derived gravity measurements at a relatively low resolution. There are many remote ocean areas in the southern hemisphere that will not be completely mapped at 400 m resolution during this decade. This study is focused on the development of synthetic bathymetry to fill the gaps. There are two types of seafloor features that are not typically well resolved by satellite gravity: abyssal hills and small seamounts (< 2.5 km tall). We generate synthetic realizations of abyssal hills by combining the measured statistical properties of mapped abyssal hills with regional geology including fossil spreading rate/orientation, rms height from satellite gravity, and sediment thickness. With recent improvements in accuracy and resolution, It is now possible to detect all seamounts taller than about 800 m in satellite-derived gravity and their location can be determined to an accuracy of better than 1 km. However, the width of the gravity anomaly is much greater than the actual width of the seamount so the seamount predicted from gravity will underestimate the true seamount height and overestimate its base dimension. In this study we use the amplitude of the vertical gravity gradient (VGG) to estimate the mass of the seamount and then use their characteristic shape, based on well surveyed seamounts, to replace the smooth predicted seamount with a seamount having a more realistic shape. </p> <p>SYNBATH_V1.2 September 20, 2021</p> <p>This version of SYNBATH has abyssal hills as described below. Superimposed on that are 30,000 gaussian seamounts with sigma to height ratios of 2.4. The heights were determined by fitting a uncompensated model VGG for a seamount of a particular height to the observed VGG in a 33 by 33 km area using a density of 2800 kg m^-3. Any seamount taller than 2600 m or less than 700 m was not used.</p> <p>SYNBATH_V1.1 July 6, 2021</p> <p>A refined version of the SYNBAPS with better blending</p> <p>SYNBATH_V1.0 July 1, 2021</p> <p>This is the first version of SYNthetic BATHymetry (SYNBATH) that is a merge of the latest SRTM15 global bathymetry/topography grid and synthetic abyssal hill fabric based on an anisotropic power spectral model published by Goff and others [2010, 2020]. The synthetic abyssal fabric fills the voids in the real bathymetry coverage that used to be filled by predicted depth.</p> <p>These are global grids with 86400 columns and 43200 rows in NETCDF format.</p> <p>Seamount Heights used in SYNBATH_V1.2 December 15, 2021</p> <p>This directory contains the locations and heights of the seamounts in the combined New and Kim Wessel (KW) catalogues. There are three categories of seamounts.</p> <p>1) good.nxybh - contains 34295 with heights successfully modeled using the VGG as described in the Sandwell 2022 publication. The file has 5 columns:</p> <p>name longitude latitude base_depth height_VGG<br> KW-00001 0.191666666667 -6.44166666667 -4060.15673828 2600<br> KW-00002 -0.425 -6.84166666667 -4125.54345703 2600<br> KW-00003 -0.075 -6.875 -4221.48730469 2600<br> .<br> .<br> .</p> <p><br> 2) uncharted.nxybh - contains 19732 seamounts that are more than 3 km from a depth sounding. The file has 5 columns:</p> <p>name longitude latitude base_depth height_VGG<br> New-00001 3.60833333333 2.74166666667 -3990.33374023 2000<br> New-00002 3.375 2.59166666667 -4113.46435547 1500<br> New-00003 3.19166666667 2.475 -4209.29638672 1200<br> .<br> .<br> .</p> <p>3) well_charted.nxybh - contains 739 seamounts that are well charted by more than 50% sounding coverage over the seamount and good coverage at the summit so the summit depth is known. The file has 6 columns:</p> <p>name longitude latitude base_depth height_VGG summit_depth<br> New-00707 -8.525 71.4916666667 -2094.12524414 1000 -1044.107788<br> New-00786 -4.775 70.0083333333 -2978.39868164 1200 -2427.73095683<br> New-00808 -4.375 66.2583333333 -3378.35644531 1100 -2656.7897947<br> .<br> .<br> .</p> <p>In addition, there are three matching kmz-files so the locations of the seamounts can be viewed in Google Earth.<br> good.kmz - yellow dots<br> uncharted.kmz - red dots<br> well_charted.kmz - green dots</p> <p> </p> <p> </p>
Model-informed target product profiles of long-acting- injectables for use as seasonal malaria prevention: code and simulation data
<p>This simulation data set and code reproduces the Figures and analysis of PLOS Global Public Health peer-reviewed article </p> <p><strong>Model-informed target product profiles of long-acting-injectables for use as seasonal malaria prevention</strong></p> <p>Authors:</p> <p>Lydia Burgert<sup>1, 2</sup>, Theresa Reiker<sup>1, 2</sup>, Monica Golumbeanu<sup>1,2</sup>, Jörg J. Möhrle<sup>1, 2, 3</sup>, Melissa A. Penny*<sup>1, 2</sup></p> <p> </p> <p><sup>1</sup> Swiss Tropical and Public Health Institute, Basel, Switzerland</p> <p><sup>2</sup> University of Basel, Basel, Switzerland</p> <p><sup>3 </sup>Medicines for Malaria Venture, Geneva, Switzerland</p> <p>*Corresponding author: <a href="mailto:melissa.penny@unibas.ch">melissa.penny@unibas.ch</a></p>
Supplementary Information: Exoplanet atmosphere retrievals in 3D using phase curve data with ARCiS: application to WASP-43b. Chubb and Min, A&A (2022).
<p>Supplementary information containing additional figures of the journal article 'Exoplanet atmosphere retrievals in 3D using phase curve data with ARCiS: application to WASP-43b' by K. L. Chubb and M. Min, published in Astronomy & Astrophysics (2022).</p>
Social information use about novel aposematic prey depends on the intensity of the observed cue
<p>Animals gather social information by observing the behavior of others, but how the intensity of observed cues influences decision-making is rarely investigated. This is crucial for understanding how social information influences ecological and evolutionary dynamics. For example, observing a predator's distaste of unpalatable prey can reduce predation by naïve birds, and help explain the evolution and maintenance of aposematic warning signals. However, previous studies have only used demonstrators that responded vigorously, showing intense beak-wiping after tasting prey. Therefore, here we conducted an experiment with blue tits (<em>Cyanistes caeruleus</em>) informed by variation in predator responses. First, we found that the response to unpalatable food varies greatly, with only few individuals performing intensive beak-wiping. We then tested how the intensity of beak-wiping influences observers' foraging choices using video-playback of a conspecific tasting a novel conspicuous prey item. Observers were provided social information from: (1) no distaste response, (2) a weak distaste response, or (3) a strong distaste response, and were then allowed to forage on evolutionarily novel (artificial) prey. Consistent with previous studies, we found that birds consumed fewer aposematic prey after seeing a strong distaste response, however a weak response did not influence foraging choices. Our results suggest that while beak-wiping is a salient cue, its information content may vary with cue intensity. Furthermore, the number of potential demonstrators in the predator population might be lower than previously thought, although determining how this influences social transmission of avoidance in the wild will require uncovering the effects of intermediate cue salience.</p>
Dataset for paper titled: Conceptual preferences can be transmitted via selective social information use between competing wild bird species
<p><a name="_Hlk66629591"></a><span>Concept learning is considered a high-level adaptive ability. Thus far, it has been studied in laboratory via asocial trial and error learning. Yet, social information use is common among animals but it remains unknown whether concept learning by observing others occurs. We tested whether pied flycatchers (</span><em><span>Ficedula hypoleuca</span></em><span>) form conceptual relationships from the apparent choices of nest-site characteristics (geometric symbol attached to the nest box) of great tits (</span><em><span>Parus major</span></em><span>). Each wild flycatcher female (n = 124) observed one tit pair that exhibited an apparent preference for either a large or a small symbol and was then allowed to choose between two nest boxes with a large and a small symbol, but the symbol shape was different to that on the tit nest. Older flycatcher females were more likely to copy the symbol size preference of tits than yearling flycatcher females when there was a high number of visible eggs or a few partially visible eggs in the tit nest. However, this depended on the phenotype; copying switched to rejection as a function of increasing body size. Possibly the quality of and overlap in resource use with the tits affected flycatchers' decisions. Hence, our results suggest that conceptual preferences can be horizontally transmitted across co-existing animals, which may increase the performance of individuals that utilize concept learning abilities in their decision-making.</span></p>
Supporting Information for Article "MarINvaders: A web toolkit of marine species for use in environmental assessments"
<p>An Excel file with the overview of number of species per ecoregion and classes of alien species is available online. The code for querying and harmonizing the databases was published as a separate package (see Lonka et al. (2021))and available with an open source (GPL v3) license at <a href="https://gitlab.com/marinvaders/marinvaders">https://gitlab.com/marinvaders/marinvaders</a>.</p>
Constraining Andean Propagation of Exhumation at the Limit of the Eastern Cordillera, NW Argentina, using Low-Temperature Thermochronology in a Structural Context - Supporting Information
<p>Supporting information accompanying the publication "Constraining Andean Propagation at the Limit of the Eastern Cordillera, NW Argentina, using Low-Temperature Thermochronology in a Structural Context" published in Tectonics. The dataset contains apatite and zircon (U-Th-Sm)/He and apatite fission track data from the Tilcara Range and San Lucas block, Jujuy, Argentina, as well as additional QTQt thermal models that are discussed in the paper.</p> <p>Table S1 contains full single-grain results from apatite fission track, apatite (AHe) (U-Th-Sm)/He and zircon (ZHe) (U-Th-Sm)/He analyses. Outliers are marked in grey and are not included in the weighted mean age. Figure S1 supports (U-Th-Sm)/He data graphically. Apatite fission track (AFT) data is supported by radial plots in Figure S2. Figure S3 shows QTQt thermal models using either AHe, AFT or ZHe single-grain ages. All of the models results are explained in the main text.</p>
Supplementary information for "Reassessment of French breeding bird population sizes using citizen science and accounting for species detectability"
<p>Reproducibility data for the manuscript "<em>Reassessment of French breeding bird population sizes using citizen science and accounting for species detectability</em>", it contains data and script for :</p> <ol> <li> <p>The R script <code>01_HDSfreq_Calibration.R</code> of the developed approach to estimate national breeding bird population size using Hierarchical Distance Sampling (HDS) and the secondary candidate set model selection method (Morin et al., 2020)</p> </li> <li> <p>The R script <code>02_pglmm_figures.R</code> for the calibration of the Phylogenetic Generalised Mixed Model (PGLMM) used in the manuscript to compare previous population size estimates to ones modelled using <code>01_HDSfreq_Calibration.R</code>, while accounting for species phylogenetic relatedness</p> </li> <li> <p>The R script <code>03_results_tables.R</code>, used to generate supplementary tables S2.1-3 and S6.1-2.</p> </li> </ol> <ul> <li> <p>Column names are highlighted in italics.</p> </li> </ul> <h2>Data description</h2> <h5>A. BirdPhylo_Burleigh_et_al.tre</h5> <p>A phylogenetic tree from Burleigh et al., 2015. Phylogenetic distances are used as random effect for the PGLMM in script <code>02_pglmm_figures.R</code></p> <h5>B. Conservation_status.txt</h5> <p>A <code>.txt</code> file of the conservation status for France (<em>Statut_FR</em>) and Europe (<em>Statut_EU</em>) for the studied species retrieved from (UICN France et al., 2016). Only <em>Statut_FR</em> is used for the table S6.1.</p> <h5>C. FBBS_trends_20122023.txt</h5> <p>A <code>.txt</code> file containing species trend of the French Breeding Bird Survey data from 2012 to 2023.</p> <ul> <li> <p>Species names (English, French) associated with FBBS trend estimated using data collected from 2012 - 2023</p> </li> <li> <p><em>Hab_specialization</em>, determined from Julliard et al. 2006 approach</p> </li> <li> <p><em>infPrec, supPerc, estimate, se, pval</em> : Species trends over 2012-2023 period in % | lower and upper confidence intervals, mean, standard error and significance</p> </li> </ul> <h5>D. PrepData_HDS.RData</h5> <p>A file containing <code>.RData</code> environment required to run <code>01_HDSfreq_Calibration.R</code> script, it contains :</p> <ul> <li> <p><strong>ATLAS12</strong> : A dataframe with breeding status information from 2012 breeding bird atlas (used to restrict model prediction grid, in regard of 2012 known breeding locations)</p> </li> <li> <p><strong>ConcordTBL</strong> : A concordance table for species names (English, French and scientific notation)</p> </li> <li> <p><strong>EPOC_ODF</strong> : observation dataset, each line corresponds to detected individuals</p> <ul> <li> <p><em>UUID, Ref, ID_liste, ID, ID_place, Grid_10x10</em> : Columns used to identify observations, lists, sites, locations, 10x10 grids</p> </li> <li> <p><em>ID_species_Biolovision, Nom_espece, english_name, scientific_name</em> : Species ID and names</p> </li> <li> <p><em>Date, Day, Month, Year, Julian_date, Obs_hour, Hour_list, Complete_checklist, Commentary, Project_name, Scheme, Observer, List_time, List_diversity, List_abundance</em> : Lists and Observation related effort covariates and metadata</p> </li> <li> <p><em>X_Lambert93_m, Y_Lambert93_m</em> : Observation locations in <code>(crs = 2154)</code></p> </li> <li> <p><em>GPS_loc_observer</em> : Logical, TRUE : location of observers corresponds to true GPS information ; FALSE : observer's location approximated as the barycenter of observations</p> </li> <li> <p><em>X_barycentre_L93, Y_barycentre_L93</em> : Observers location in <code>(crs = 2154)</code></p> </li> <li> <p><em>Use_distance_sampling, Observation_distance_m, Distance_bin_logical, Distance_class_0_25, Distance_class_25_100, Distance_class_100_200, Distance_class_200_more</em> : Distance sampling related informations</p> </li> <li> <p><em>Abudance_brut, Estimate, Number, Nb_male_identified, Nb_female_identified, Nb_juvenile_identified, Nb_grounded, Nb_flying, Nb_auditory, Nb_NA</em> : Observation metadata, used in case of <em>a priori</em> filter over male detection.</p> </li> </ul> </li> <li> <p><strong>grid_pred_envvar</strong> : Prediction grid with environmental covariates, see appendix S3 of the manuscript, covering metropolitan France</p> </li> <li> <p><strong>grid_pred.sf</strong> : corresponding sf object</p> </li> <li> <p><strong>L93_10x10</strong> : sf object corresponding to 10x10 grid used in 2012 atlas</p> </li> <li> <p><strong>ObsVar_EPOCODF</strong> : dataframe specifying lists effort covariates</p> </li> <li> <p><strong>OCCU_EPOC_ODF</strong> : Environmental covariate agregated over lists</p> </li> <li> <p><strong>OCCU_EPOC_ODF_sites_envvar</strong> : Environmental covariate agregated over sites</p> </li> <li> <p><strong>table.pheno</strong> : species table specifying related phenology filter</p> </li> </ul> <h5>E. ReadOutput_HDSfreq_comparison.csv</h5> <p>A <code>.csv</code> table of species population size estimated using 2021-2023 EPOC-ODF data over areas determined as breeding in the 2012 atlas. <strong>Predictions were restrained over location known as breeding in 2012 for the sake of comparison.</strong></p> <ul> <li> <p><em>HDS_estimUnfenced_XXX</em> : average pop. size estimated with confidence interval before prediction post-treatment (describe in fig 2. of the manuscript)</p> </li> <li> <p><em>HDS_estim_ExtrapolFence_XXX</em> : average pop. size estimated with confidence interval after prediction post-treatment</p> </li> <li> <p><em>NB_data_calib</em> : Number of observations (distance data, not sites) used for calibration</p> </li> <li> <p><em>MALE_FILTERING</em> : (logical) indicating if female individuals could be detected in the same proportion of males during list recording. FALSE : we considered that estimated pop.size corresponded to the number of individuals leading to a division by 2 for the comparison with the previous atlas (in pairs). (cf . line 88-90 in <code>02_pglmm_figures.R</code>)</p> </li> <li> <p><em>EcartDTF_filtrage_maleOnly</em> : If MALE_FILTERING == T, proportion of the remaining data used for calibration after removal of list with individual tagged as female/juvenile (in %)</p> </li> <li> <p><em>Max_dist_breaks</em> : Maximal distance for detection function, after right-side truncation of 5%</p> </li> <li> <p><em>Chat</em> : Coefficient of overdisperion of the best model in the second candidate set</p> </li> <li> <p><em>MED_MEAN_Prob_Detect</em> (.._SE) : weighted averaged median of intercept from the availability state from HDS models, weigthed AICc-wise</p> </li> <li> <p><em>MED_MEAN_Density</em> : weighted averaged median of intercept from the abundance state from HDS models, weigthed AICc-wise</p> </li> <li> <p><em>Significant_phi/lambda</em> : Categorial (Significant/Near/Not), are availability/abundance intercepts significatively different from 0 (significant : alpha = 0.05, near : alpha = 0.1)</p> </li> <li> <p><em>KeyFun_used</em> : Key function used for distance sampling</p> </li> <li> <p><em>Mixtured_used</em> : Mixture used in the abundance state for HDS</p> </li> </ul> <h5>F. ReadOutput_HDSfreq_comparison_20212022.csv</h5> <p>A <code>.csv</code> table of species population size estimated using 2021-2022 EPOC-ODF data over areas determined as breeding in the 2012 atlas. Used for the robustness analysis of HDS estimated population size, see appendix S2 and table S2.2 of the manuscript.</p> <h5>G. ReadOutput_HDSfreq_EstimMetropole.csv</h5> <p>A <code>.csv</code> table of species population size estimated using 2021-2023 EPOC-ODF data over <strong>metropolitan France</strong>.</p> <ul> <li> <p><em>HDS_estimUnfenced_XXX</em> : average pop. size estimated with confidence interval before prediction post-treatment (describe in fig 2. of the manuscript)</p> </li> <li> <p><em>HDS_estim_ExtrapolFence_XXX</em> : average pop. size estimated with confidence interval after prediction post-treatment</p> </li> <li> <p><em>MALE_FILTERING</em> : (logical) indicating if female individuals could be detected in the same proportion of males during list recording. FALSE : we considered that estimated pop.size corresponded to the number of individuals leading to a division by 2 for conversion to pop. size in breeding pairs</p> </li> <li> <p><em>Chat</em> : Coefficient of overdisperion of the best model in the second candidate set</p> </li> <li> <p><em>KeyFun_used</em> : Key function used for distance sampling</p> </li> <li> <p><em>Mixtured_used</em> : Mixture used in the abundance state for HDS</p> </li> </ul> <h5>H. TABLE_SpeciesFilters_and_2012Estimates.txt</h5> <p>A <code>.txt </code>table containing species names (English, French and scientific notation), filters and 2012 French atlas pop. size estimates</p> <ul> <li> <p><em>debut_jour</em> : starting day of the month for phenology filter</p> </li> <li> <p><em>debut_mois</em> : starting month for phenology filter</p> </li> <li> <p><em>fin_jour</em> : ending day of the month for phenology filter</p> </li> <li> <p><em>fin_mois</em> : ending month for phenology filter</p> </li> <li> <p><em>Estim_low/up_Atlas2012</em> : Lower and Upper interval of estimated pop. size in 2012 (number in breeding pairs)</p> </li> <li> <p><em>gregarious</em> : logical (0,1) specifying if the species is considered gregarious during its breeding season</p> </li> </ul> <h5>I. sessionInfo_script_XX</h5> <p>User R session information, obtained from <code>sessionInfo()</code> R function, used for running R script.</p> <h2>Code</h2> <h5>A. <code>01_HDSfreq_Calibration.R</code></h5> <p>R script showcasing data formatting and model calibration of the HDS based upon frequentist aproach from <code>unmarked</code> R package. For more details of the model calibration approach, see appendix S4 of the manuscript.</p> <h5>B. <code>02_pglmm_figures.R</code></h5> <p>Script for the calibration of the PGLMM and generation of figure 5 of the manuscript.</p> <h5>C. <code>03_results_tables.R</code></h5> <p>Script to generate tables depicted in Appendices S2 (S2.1-3) and S6 (S6.1-2)</p> <h5>D. <code>HDS_functions.R</code></h5> <p>R script called in <code>01_HDSfreq_Calibration.R</code>, contains 2 functions:</p> <ul> <li> <p><code>Try_HDS()</code> : Function implementing a try-catch permitting calibration of multiple species in a loop.</p> <ul> <li> <p>Species with non convergent models are skipped sending a notification to the user R interface.</p> </li> <li> <p>Used in all sub-candidate sets (i.e. "null", "p", "phi", "lambda")</p> </li> <li> <p>When phase="ALL" corresponding to the second candidate set (i.e. ensemble of best model candidates, with delta_AIC <= 10, from previous sub-candidate sets), it permits the use of previous sub-candidates set coefficients as starting values, with <code>StartValues </code>argument</p> </li> <li> <p>Later part of the function hack the call of the unmarkedFit class, in order to accommodate from calibrating a gdistsamp using characters formulas</p> </li> </ul> </li> <li> <p><code>fitstats()</code> : Function from unmarked::parboot(), available with <code>help(parboot)</code>. Allow estimation of multiple goodness-of-git statistic (Freeman-Tukey, Chi-squared and Sum of Squared Estimate of errors) through parametric bootstrap. In the manuscript, only chi-squared metric is used.</p> </li> </ul> <h5>E. <code>dsmextra_modif_function.R</code></h5> <p>R script called in <code>01_HDSfreq_Calibration.R</code>. Miscellaneous adjustment of core function from <code>dsmextra </code>package (main change being the integration of tolerance argument (<code>tol</code>) in the chain of function.</p> <h5>F. <code>misc_unmarked.R</code></h5> <p>R script called in <code>01_HDSfreq_Calibration.R</code>. modify Setmethods for unmarked function, in particular for <code>unmarked::parboot</code>, allowing parallelization of parametric bootstrap with prior unmarked version (<code>unmarked < 1.3.0</code>).</p> <h2>References</h2> <p>Data was derived from the following sources:</p> <ul> <li> <p>Burleigh, J.G., Kimball, R.T., Braun, E.L., 2015. Building the avian tree of life using a large-scale, sparse supermatrix. Molecular Phylogenetics and Evolution 84, 53–63. <a href="https://doi.org/10.1016/j.ympev.2014.12.003">https://doi.org/10.1016/j.ympev.2014.12.003</a></p> </li> </ul> <p>Other sources :</p> <ul> <li> <p>Julliard, R., Clavel, J., Devictor, V., Jiguet, F., Couvet, D., 2006. Spatial segregation of specialists and generalists in bird communities. Ecology Letters 9, 1237–1244. <a href="https://doi.org/10.1111/j.1461-0248.2006.00977.x">https://doi.org/10.1111/j.1461-0248.2006.00977.x</a></p> </li> <li> <p>Morin, D.J., Yackulic, C.B., Diffendorfer, J.E., Lesmeister, D.B., Nielsen, C.K., Reid, J., Schauber, E.M., 2020. Is your ad hoc model selection strategy affecting your multimodel inference? Ecosphere 11, e02997. <a href="https://doi.org/10.1002/ecs2.2997">https://doi.org/10.1002/ecs2.2997</a></p> </li> <li> <p>UICN France, MNHN, LPO, SEOF, ONCFS, 2016. La Liste rouge des espèces menacées en France - Chapitre Oiseaux de France métropolitaine. Paris, France.</p> </li> </ul>
OccuTherm: Occupant Thermal Comfort Inference using Body Shape Information
<p><strong>OccuTherm: Occupant Thermal Comfort Inference using Body Shape Information</strong></p> <p>This repository contains the official data from a USDOE-funded project at Carnegie Mellon University and Bosch Research Pittsburgh. </p> <p>The primary goal of the project was to investigate the relationship between indoor commercial building occupant thermal comfort and various biometric and environmental predictors. We performed 77 individual comfort experiments, approved by our Institutional Review Board (IRB) and in satisfaction of participant consent guidelines. Our goal was to generate a dataset than enables comprehensive study of human thermal comfort preferences, in a commercial building environment, across a wide range of indoor environmental conditions. The data is comprised of the following feature groups: depth camera frames, biometrics sensor data, body shape information, subjective comfort data from the mobile device application, environmental sensor data from the commercial building HVAC system, and outdoor weather station data.</p> <p>This is the official dataset release for the following conference paper:</p> <blockquote> <p>Jonathan Francis*, Matias Quintana*, Nadine von Frankenberg, Sirajum Munir, and Mario Bergés. 2019. OccuTherm: Occupant Thermal Comfort Inference using Body Shape Information. In BuildSys '19: ACM International Conference on Systems for Energy-Efficient Buildings, Cities, and Transportation, November 13–14, 2019, New York, NY. ACM, New York, NY, USA, 10 pages.</p> </blockquote> <p>To use this dataset, first download <strong>all</strong> the files.</p> <p>Next, issue the following commands on, e.g., Linux terminal:</p> <pre><code class="language-bash">>$ cd /path/to/dataset/files >$ cat occutherm_dataset_v0-0-0.tar.gza* > archive.tar.gz >$ tar -xvzf archive.tar.gz</code></pre> <p>Modeling and mobile application code are available in our project repository: <a href="https://github.com/jonfranc/occutherm">https://github.com/jonfranc/occutherm</a></p> <p>If you find the repository or the dataset useful, please cite our paper:</p> <pre><code>@inproceedings{francis_buildsys2019, author = {Francis, Jonathan and Quintana, Matias and von Frankenberg, Nadine and Munir, Sirajum and Berges, Mario}, title = {OccuTherm: Occupant Thermal Comfort Inference using Body Shape Information}, booktitle = {Proceedings of the 6th International Conference on Systems for Energy-Efficient Buildings, Cities, and Transportation}, series = {BuildSys '19}, year = {2019}, isbn = {978-1-4503-7005-9/19/11}, location = {New York, NY}, numpages = {10}, acmid = {3360858}, publisher = {ACM}, address = {New York, NY, USA}, keywords = {Thermal Comfort, Human Studies, Machine Learning}, }</code></pre> <p> </p>
Data for "Using physics-informed neural networks to predict the lifetime of laser powder bed fusion processed 316L stainless steel under multiaxial low-cycle fatigue loading"
<p>Title of dataset: Data for "Using physics-informed neural networks to predict the lifetime of laser powder bed fusion processed 316L stainless steel under multiaxial low-cycle fatigue loading".</p> <p>Name/institution/contact information: Dr. Michal Bartošák, Czech Technical University in Prague - Faculty of Mechanical Engineering, email: michal.bartosak@fs.cvut.cz.</p> <p>Date of data collection: The data were collected between 2021 and 2024.</p> <p>File name structure: The data consists of two files: "316L_fatigue_and_defects.xls," which contains fatigue lifetime data and defect characteristics, and an associated description file, "read_me.txt."</p> <p>See "https://doi.org/10.1016/j.ijfatigue.2024.108608" for the associated article and a detailed description of the methods.</p>
Using interview surveys and multispecies occupancy models to inform vertebrate conservation
<p>The Excel workbook WGAllSpeciesCov.xlsx contains detection histories of 30 species of vertebrates in 395 sites in the Western Ghats, India obtained from interviews with field staff of the Forest Department, people from local communities and formally trained people. The workbook also contains site level covariates or determinants of species occurrence. The associated README text file contains metadata to describe the data in the xlsx workbook. Complete R code to read in the data and run the model is provided in the R file: FP_MSp_SS_SpR_AllTraitGr.R and the JAGS code for the multispecies occupancy model is provided in the text file FP_MSp_SS_SpR_AllTraitGr_Exp_NoFP.txt</p>
Relations from Italian Wikipedia using Unsupervised Information Extraction
<p>This dataset contains relations extracted from the Italian Wikipedia by the WikiOIE framework.<br> WikiOIE is based on UDPipe and the Universal Dependencies project for text processing.<br> It easily allows customizing the information extraction (IE) approach to automatically extract triples (subject, predicate, object).<br> This dataset contains relations extracted by two unsupervised IE methods. The former (<strong>simple</strong>) is based only on PoS-tag patterns; the latter (<strong>simpledep</strong>) also uses syntactic dependencies. <br> The extraction process is provided in JSON format.</p> <p>More information and the Java code are available here https://github.com/pippokill/WikiOIE</p> <p>Pierluigi Cassotti, Lucia Siciliani, Pierpaolo Basile,Marco de Gemmis, and Pasquale Lops. 2021. Extracting relations from Italian Wikipedia using unsupervised information extraction. In Proceedings of the 11th Italian Information Retrieval Workshop 2021 (IIR 2021). CEUR-WS.</p>
PB preprocessed data used in paper "Multi variables time series information bottleneck"
<p>Preprocessed PB data used in paper "Multi variables time series information bottleneck" with the <a href="https://github.com/DenisUllmann/IB-MTS">GitHub</a> code</p> <p>This dataset is created from a public available dataset of solar power data collected in Alabama by <a href="https://www.nrel.gov/grid/solar-power-data.html">C</a><a href="https://pems.dot.ca.gov/">alTrans</a>.</p> <p>The npz file is a numpy (np) compressed data and can be loaded using np.load with allow_pickle=True<br> Loaded data is then a python dict described bellow.</p> <p>Each sample 'data' is a np.ndarray with 2 dimensions: time (various length) and wavelength (length=325 representing 325 traffic detectors ordered like in <a href="https://www.nrel.gov/grid/solar-power-data.html">C</a><a href="https://pems.dot.ca.gov/">alTrans</a>).</p> <p>Each sample is given a 'position' which is a list of length 4:<br> position[1] is a string that gives the name of the event<br> position[4] is a boolean vector that gives the time positionsof the corresponding sample in the original sequence of public IRIS level2 data</p> <p>Data file info :<br> Type: .npz<br> Size: 114.23MB<br> *** Key: 'data_TR_PB'<br> ndarray data of length 3<br> containing np.ndarray of shapes [12160, 325]</p> <p>*** Key: 'data_VAL_PB'<br> ndarray data of length 3<br> containing np.ndarray of shapes [868, 325]</p> <p>*** Key: 'data_TE_PB'<br> ndarray data of length 3<br> containing np.ndarray of shapes [4343, 325]</p> <p>*** Key: 'data_TR'<br> ndarray data of length 3<br> containing np.ndarray of shapes [12160, 325]</p> <p>*** Key: 'data_VAL'<br> ndarray data of length 3<br> containing np.ndarray of shapes [868, 325]</p> <p>*** Key: 'data_TE'<br> ndarray data of length 3<br> containing np.ndarray of shapes [4343, 325]</p> <p>*** Key: 'position_TR_PB'<br> ndarray data of length 3<br> containing ndarray data of length 4<br> containing mix of types {'str', 'ndarray', 'int'}</p> <p>*** Key: 'position_VAL_PB'<br> ndarray data of length 3<br> containing ndarray data of length 4<br> containing mix of types {'str', 'ndarray', 'int'}</p> <p>*** Key: 'position_TE_PB'<br> ndarray data of length 3<br> containing ndarray data of length 4<br> containing mix of types {'str', 'ndarray', 'int'}</p> <p>*** Key: 'position_TR'<br> ndarray data of length 3<br> containing ndarray data of length 4<br> containing mix of types {'str', 'ndarray', 'int'}</p> <p>*** Key: 'position_VAL'<br> ndarray data of length 3<br> containing ndarray data of length 4<br> containing mix of types {'str', 'ndarray', 'int'}</p> <p>*** Key: 'position_TE'<br> ndarray data of length 3<br> containing ndarray data of length 4<br> containing mix of types {'str', 'ndarray', 'int'}</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.