Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9,786
datasets available to search
ShareScore release 0.7.1
Dataset results
9,786 results for “selection”
Supplementary data for "Influence of prey availability on habitat selection during the non-breeding period in a resident bird of prey"
<p><strong>Abstract</strong></p> <p>Background: For resident birds of prey in the temperate zone, the cold non-breeding period can have strong impacts on survival and reproduction with implications for population dynamics. Therefore, the non-breeding period should receive the same attention as other parts of the annual life cycle. Birds of prey in intensively managed agricultural areas are repeatedly confronted with unpredictable, rapid changes in their habitat due to agricultural practices such as mowing, harvesting, and ploughing. Such a dynamic landscape likely affects prey distribution and availability and may even result in changes in habitat selection of the predator throughout the annual cycle.</p> <p>Methods: In the present study, we 1) quantified barn owl prey availability in different habitats across the annual cycle, 2) quantified the size and location of barn owl breeding and non-breeding home ranges using GPS-data, 3) assessed habitat selection in relation to prey availability during the non-breeding period, and 4) discussed differences in habitat selection during the non-breeding period to habitat selection during the breeding period.</p> <p>Results: The patchier prey distribution during the non-breeding period compared to the breeding period led to habitat selection towards grassland during the non-breeding period. The size of barn owl home ranges during breeding and non-breeding were similar, but there was a small shift in home range location which was more pronounced in females than males. The changes in prey availability led to a mainly grassland-oriented habitat selection during the non-breeding period. Further, our results showed the importance of biodiversity promotion areas and undisturbed field margins within the intensively managed agricultural landscape. </p> <p>Conclusions: We showed that different prey availability in habitat categories can lead to changes in habitat preference between the breeding and the non-breeding period. Given these results we show how important it is to maintain and enhance structural diversity in intensive agricultural landscapes, to effectively protect birds of prey specialised on small mammals. Hereafter we provide the datasets and R script to reproduce the resource selection functions.</p>
Heterogeneous Habenular Neuronal Ensembles during Selection of Defensive Behaviors
<p>Optimal selection of threat-driven defensive behaviors is paramount to an animal's survival. The lateral habenula (LHb) is a key neuronal hub coordinating behavioral responses to aversive stimuli. Yet, how individual LHb neurons represent defensive behaviors in response to threats remains unknown. Here, we show that in mice, a visual threat promotes distinct defensive behaviors, namely runaway (escape) and action-locking (immobile-like). Fiber photometry of bulk LHb neuronal activity in behaving animals reveals an increase and a decrease in calcium signal time-locked with runaway and action-locking, respectively. Imaging single-cell calcium dynamics across distinct threat-driven behaviors identify independently active LHb neuronal clusters. These clusters participate during specific time epochs of defensive behaviors. Decoding analysis of this neuronal activity reveals that some LHb clusters either predict the upcoming selection of the defensive action or represent the selected action. Thus, heterogeneous neuronal clusters in LHb predict or reflect the selection of distinct threat-driven defensive behaviors.</p>
Rereferenced Chemical Shift files used to select dihedrals for the 5 secondary structure conformations used to calculate Conformational Variability (ConVa).
<p>These Chemical Shifts in this repository were processed with ShiftCrypt (1) and then processed as described in the manuscript. </p> <p> </p> <p>1- Gabriele Orlando, Daniele Raimondi, Luciano Porto Kagami, Wim F Vranken, ShiftCrypt: a web server to understand and biophysically align proteins through their NMR chemical shift values, <em>Nucleic Acids Research</em>, Volume 48, Issue W1, 02 July 2020, Pages W36–W40, <a href="https://doi.org/10.1093/nar/gkaa391">https://doi.org/10.1093/nar/gkaa391</a></p>
Data for: Detecting Long-Term Balancing Selection Using Allele Frequency Correlation
<p>Genome-wide and top 1% scores for 1KG project data output from BetaScan reported in:</p> <p><a href="https://pubmed.ncbi.nlm.nih.gov/28981714/">Detecting Long-Term Balancing Selection Using Allele Frequency Correlation.</a></p> <p>Siewert KM, Voight BF. Mol Biol Evol. 2017 Nov 1;34(11):2996-3005. doi: 10.1093/molbev/msx209.</p> <p>PMID: 28981714</p> <p>Code available at: https://github.com/ksiewert/BetaScan</p>
Data for: BetaScan2: Standardized Statistics to Detect Balancing Selection Utilizing Substitution Data
<p>Genome-wide scan using BetaScan2 in 1KG populations report in:</p> <p><a href="https://pubmed.ncbi.nlm.nih.gov/32011695/">BetaScan2: Standardized Statistics to Detect Balancing Selection Utilizing Substitution Data.</a></p> <p>Siewert KM, Voight BF.Genome Biol Evol. 2020 Feb 1;12(2):3873-3877. doi: 10.1093/gbe/evaa013.</p> <p>PMID: 32011695 </p> <p>Code available at: https://github.com/ksiewert/BetaScan</p>
Data for: Patterns of shared signatures of recent positive selection across human populations
<p>Genome-wide summary stats for modified iHS scan in 1KG as reported in:</p> <p><a href="https://pubmed.ncbi.nlm.nih.gov/29459708/">Patterns of shared signatures of recent positive selection across human populations.</a></p> <p>Johnson KE, Voight BF.Nat Ecol Evol. 2018 Apr;2(4):713-720. doi: 10.1038/s41559-018-0478-6. Epub 2018 Feb 19.</p> <p>PMID: 29459708</p> <p>Code available at: https://github.com/bvoightlab/iHS_calc</p>
Selective Attention VR and PC : Data And Analysis
<p><strong>Data and Analysis Repository for</strong></p> <p><strong>Developing Virtual Reality and Computer Screen Experiments One to One Using Selective Attention as a Case Study</strong></p> <p>June 2023, Rasmus Ahmt Hansen and Marta Topor</p> <p>The current repository holds all data and analysis scripts used in the report named above. Data files are saved in .csv format and analysis scripts were written using R and R Markdown.</p> <p>The report preprint can be accessed at:</p> <p>The study aimed to develop a reliable PC control condition for a VR experiment assessing selective attention in grade 0 children.<br> The selective attention task we developed and implemented can be accessed here:</p> <ul> <li>PC : <a href="https://doi.org/10.5281/zenodo.7844487">https://doi.org/10.5281/zenodo.7844487</a></li> <li>VR : <a href="https://doi.org/10.5281/zenodo.7844593">https://doi.org/10.5281/zenodo.7844593</a></li> </ul> <p><strong>Participants</strong></p> <p>73 grade 0 children from Danish primary schools completed the selective attention test in both VR and PC environments. Performance quality was low and thus we only included 19 participants in final analyses. All data, included and excluded, are openly available in this repository.</p> <p><strong>Data</strong></p> <ul> <li>Raw data from the PC condition can be found in the RAW PC folder</li> <li>Raw data from the VR condition can be found in the RAW VR folder</li> <li>Demographic data, anonymised, can be found in the demographics.csv file</li> <li>The final data from the 19 participants included in statistical analyses can be found in the final_data_set.csv file</li> </ul> <p><strong>Analysis</strong></p> <ul> <li>Demographic analyses can be found in the demographics.R file</li> <li>Data processing, quality control and statistical analyses can be found in the full_analysis_script.Rmd</li> <li>The plots folder holds plots used in the study report</li> </ul>
Selection of Low Frequency Extensions of Saturn Kilometric Radiation detected by Cassini/RPWS.
<p>Selection of Low Frequency Extensions (LFEs) of Saturn Kilometric Radiation (SKR) observed by Cassini RPWS. The catalogue of LFES are presented in TFCAT format (Time-Frequency Catalogue, https://gitlab.obspm.fr/maser/catalogues/tfcat/-/tree/master/). This work was funded by Science Foundation Ireland grant 18/FRL/6199. </p>
Functional potential and evolutionary response to long-term heat selection of bacterial associates of coral photosymbionts
<p>Sequencing reads were assembled using the genome assembler pipeline Shovill v1.1.0. Briefly, the Shovill pipeline included read trimming using Trimmomatic v0.39, de novo assembly with SPAdes v3.15.5 and genome polishing with Pilon v1.24. After the pipeline, additional polishing was performed by mapping the reads back to the contigs with BWA v0.7.17 and sorting the resulting SAM/BAM files using SAMtools v1.15.1. Pilon v1.24 was then used to correct bases, fix mis-assemblies and fill gaps. The reformat.sh script from the Bbmap package v38.76 (-minlength=1000) was used to filter out contigs less than 1000bp. The draft genome assemblies were then annotated with Bakta v1.7.0. </p> <p>Single nucleotide polymorphism (SNP) detection between WT (WTref) and SS (SSref) samples were then performed using snippy v4.6.0, where both WT and SS samples were inputted as the reference genome in turn.</p> <p>A subset of the snippy output files are uploaded here and contain all variants found.</p>
VESPA: Analysis of selected CPTAC datasets
<p>This is a supplemental dataset to the VESPA manuscript "<a href="https://doi.org/10.1101/2023.02.15.528736">Network-based elucidation of colon cancer drug resistance by phosphoproteomic time-series analysis</a>". VESPA was used to reconstruct signaling networks and to conduct kinase/phosphatase activity inference for selected subsets of the following CPTAC datasets:</p> <ul> <li><a href="https://pubmed.ncbi.nlm.nih.gov/31031003/">CPTAC S037 (COAD)</a></li> <li><a href="https://pubmed.ncbi.nlm.nih.gov/31675502/">CPTAC S044 (ccRCC)</a></li> <li><a href="https://pubmed.ncbi.nlm.nih.gov/32649874/">CPTAC S046 (LUAD)</a></li> <li><a href="https://pubmed.ncbi.nlm.nih.gov/33242424/">CPTAC S047 (PBT)</a></li> <li><a href="https://pubmed.ncbi.nlm.nih.gov/31585088/">CPTAC S049 (HBV-HCC)</a></li> <li><a href="https://pubmed.ncbi.nlm.nih.gov/34358469/">CPTAC S058 (LSCC)</a></li> <li><a href="https://pubmed.ncbi.nlm.nih.gov/34534465/">CPTAC S061 (PDAC)</a></li> </ul>
Selected near-bottom and other variables from NW European shelf physics-biogeochemistry downscaled ocean climate projections, 3-member ensemble.
<p>Selected fields of physical and biogeochemical ocean variables from a 3-member ensemble of coupled physics-biogeochemistry downscaled climate runs on the North Western European Continental Shelf. All ensemble members use the NEMO-ERSEM model suite and cover the 1990-2099 period. Easch member is foced with a different set of atmospheric and oceanic boundary conditions from one of three CMIP5 ESMs that are: HADGEM2-ES, IPSL-CM5A-MR and GFDL-ESM2G. This dataset contains monthly average values saved as 2D fields either near-bottom, at the surface or depth integrated. The variables here saved are near-bottom oxygen, oxygen solubility, oxygen saturation state, temperature and bacterial respiration, surface salinity, depth integrated net primary production, and potential energy anomaly. Additionally the Western Norwegian Trench Current flux is provided (its values come smoothed with a gaussian filter). reference publication: https://doi.org/10.5194/egusphere-2023-1049. The complete set of variables is available from the authors upon request.</p>
Enhanced Westermo dataset - Transformed and Modified for Test case Selection and Priorotization in the context of Continuous Integration and Reinforcement Learning.
<p><strong>Overview</strong></p> <p>This repository contains a modified version of the existing, recently published dataset, Westermo. The initial dataset was gathered at Westermo Network Technologies AB, located in Västerås, Sweden. It encompasses over <strong>1 Million verdicts</strong> obtained from testing embedded systems, collected over a span of more than <strong>500 consecutive days</strong> of nightly testing. The dataset has been transformed and tailored specifically to cater to the research community, particularly for addressing challenges such as regression test selection, identification of flaky tests, and visualization of test results. The original dataset can be accessed through the reference provided in <strong>[1]</strong>.</p> <p>The Westermo dataset offers valuable historical information regarding the execution of test cases and their corresponding results. It serves as a valuable resource for evaluating and comparing different Test case Selection and Prioritization (TSP) techniques, enabling researchers to identify test cases that are more likely to fail during subsequent executions. Test cases in the dataset are characterized by attributes such as execution duration, previous last execution time, and the results of their recent executions.</p> <p>This dataset offers valuable historical information regarding the execution of test cases and their corresponding results. It serves as a valuable resource for evaluating and comparing different test case prioritization and selection techniques, enabling researchers to identify test cases that are more likely to fail during subsequent executions. Test cases in the dataset are characterized by attributes such as execution duration, previous last execution time, and the results of their recent executions.</p> <table align="left"> <caption><strong>Table 1: Dataset Overview</strong></caption> <tbody> <tr> <td>Test Cases</td> <td>1855</td> </tr> <tr> <td>CI Cycles</td> <td>15,197</td> </tr> <tr> <td>Verdict</td> <td>1,036,818</td> </tr> <tr> <td>Failed</td> <td>5.03%</td> </tr> </tbody> </table> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p>However, the diversity and multitude of the features in the dataset can be irrelevant to some TSP approaches. This led us to perform a dataset conversion, where we customized Westermo to have the same features from Paint Control and IOF/ROL, two widely used datasets in Reinforcement Learning based TSP approaches.</p> <p>This conversion required the combination of multiple variables and generating the target ones. When it comes to generating the “LastResults” and “Cycle” values, further analysis was required and the data handling needed an in-depth understanding of how the nightly testing was conducted. This led us to investigate what a CI cycle is in their context, and we followed their definition of a session, stating that “a session is when we run a suite of tests on one test system with a certain software version and testware version”. When splitting the data according to the 9 different systems used, we were able to generate 9 different sub-sets that fit the CI context.</p> <p> </p> <p><strong>File Format</strong></p> <p>The compressed .zip file contains 9 files, each one corresponding to each of the 9 systems. The datasets are available in CSV format, with the semicolon (;) serving as the delimiter. The columns included are represented in the table below along with their descriptions.</p> <table> <caption><strong>Table 2: Parameters of the dataset</strong></caption> <thead> <tr> <th scope="col">Column Name</th> <th scope="col">Content</th> </tr> </thead> <tbody> <tr> <td>Id</td> <td>Unique numeric identifier of the test execution </td> </tr> <tr> <td>Name</td> <td>Unique numeric identifier of the test case</td> </tr> <tr> <td>Duration</td> <td>Approximated runtime of the test case</td> </tr> <tr> <td>CalcPrio</td> <td>Priority of the test case, calculated by the prioritization algorithm (output column, initially 0)</td> </tr> <tr> <td>LastRun</td> <td>Previous last execution of the test case as date-time-string (Format: <em>YYYY-MM-DD HH:ii </em>)</td> </tr> <tr> <td>LastResults</td> <td>List of previous test results (Failed: 1, Passed: 0), ordered by ascending age. Lists are delimited by [ ].</td> </tr> <tr> <td>Verdict</td> <td> <p>Test verdict of this test execution (Failed: 1, Passed: 0)</p> </td> </tr> <tr> <td>Cycle</td> <td>The number of the CI cycle this test execution belongs to.</td> </tr> </tbody> </table> <p> </p> <p>The implications of this conversion are important as it can help the previous works to re-assess their approaches and have more data for training and testing, as well as opening a broader data spectrum for future researchers in this field to find ready-to-use, rich datasets, on which they could evaluate their approaches and contribute to the TSP community. This also addresses the limitations in the field discussed in the systematic literature review <strong>[2]</strong>, stating that future research on TSP techniques should focus on collecting data from more recent subjects in a CI context with varying failure rates and larger execution times, as reproducible studies with appropriate datasets are needed to develop a usable body of knowledge regarding TSP over time. We believe that this conversion of the Westermo dataset is our contribution to alleviating the gap for the RL-based approaches.</p> <p>The original dataset can be found <a href="https://sites.mdu.se/aidoart/results/open-source/test-results-dataset-westermo">here.</a></p>
A Harmonised Dataset for Modelling Select Underutilised Crops Across EU
<p>Version 2: the data was checked and refined against issues that were found in the columns bulk density and treatments. </p> <p>Datasets are a compilation of information that are collected from various sources including books, research articles, databases, website articles, experts and local communities growing underutilised crops, etc. The focus was on extracting data from literature sources that were mainly peer reviewed, credible, and are primarily published in English language. </p> <p>Taxonomy data: crop, variety or landrace</p> <p>Publication information: author, journal, year</p> <p>Geographic data: continent, country, site, latitude, longitude</p> <p>Soil data: Lower Depth, Clay (%), Sand (%), Silt (%), Texture, S.O., Bulk Density, Total carbon (%)</p> <p>Experimental data: Experiment duration (years), Seeding rate, Sowing depth, dates of each treatment /intervention.</p> <p>Production: yield, yield date, Above Ground Biomass (Agb), Agb date</p> <p>Phenology data: Growing Degree Days (sowing-to-harvest), emergence date, flowering date, maturity date</p> <p>Metadata: Study number, DOI, link </p> <p>Credible sources containing experimental data, either from agronomy trials or meta-analysis containing experimental data were selected through literature search. As the focus of the work was to collate as much information as possible about agronomy trials of select underutilised crops, all data were collected and inserted into a shareable template on Google Docs. </p> <p>List of crops:<br> <br> <em>Ceratonia siliqua</em>, Carob, Underutilised legume tree<br> <em>Cichorium endivia</em>, Endive, Underutilised vegetable<br> <em>Eragrostis tef</em>, Teff, Underutilised cereal<br> <em>Ficus carica</em>, Fig, Underutilised fruit tree<br> <em>Helianthus tuberosus</em>, Jerusalem artichoke, Underutilised starchy roots/tubers<br> <em>Lentil Culinaris</em>, Lentil, Pulse<br> <em>Lupinus albus</em>, White lupin, Underutilised legume<br> <em>Malus domestica</em>, Apple, Fruit<br> <em>Malus pumila</em>, Apple, Fruit<br> <em>Medicago sativa</em>, Alfaalfa, Underutilised legume<br> <em>Panicum miliaceum</em> , Proso millet, Underutilised minor millet<br> <em>Pisum sativum</em>, Pea, Legume<br> <em>Prunus avium</em>, Cherry, Fruit<br> <em>Prunus domestica</em>, Plum, Fruit<br> <em>Pyrus communis,</em> Pear, Fruit<br> <em>Rheum rhaponticum</em>, Rhubarb, Underutilised vegetable<br> <em>Setaria italic</em>, Foxtail millet, Underutilised minor millet<br> <em>Trifolium repens</em> , Clover, Underutilised legume<br> <em>Vicia faba</em>, Faba bean, Underutilised legume<br> <em>Vigna unguiculata</em>, Cowpea, Underutilised legume</p>
ERA5-Land selected indicators daily aggregates for the Latin America region, 1951
<p>This deposit contains NetCDF files with daily aggregates from Copernicus Era5-Land eight selected indicators, covering the Latin America region, for 1951.</p><p>Each file represents one indicator aggregation for one month of the year. Inside each NetCDF file, the layers contain the daily aggregates.</p><p>For 2m dewpoint pressure, 10m u component of wind, 10m v component of wind, surface pressure, the mean function was used for aggregation. For total precipitation, the sum function was used for aggregation. For 2m temperature, the functions maximum, mean and minimum were used for aggregation.</p><p>Those files were created using the <a href="https://github.com/ErikKusch/KrigR">KrigR</a> package.</p>
ERA5-Land selected indicators daily aggregates for the Latin America region, 1950
<p>This deposit contains NetCDF files with daily aggregates from Copernicus Era5-Land eight selected indicators, covering the Latin America region, for 1950.</p><p>Each file represents one indicator aggregation for one month of the year. Inside each NetCDF file, the layers contain the daily aggregates.</p><p>For 2m dewpoint pressure, 10m u component of wind, 10m v component of wind, surface pressure, the mean function was used for aggregation. For total precipitation, the sum function was used for aggregation. For 2m temperature, the functions maximun, mean and minimum were used for aggregation.</p><p>Those files were created using the <a href="https://github.com/ErikKusch/KrigR">KrigR</a> package.</p>
ERA5-Land selected indicators daily aggregates for the Latin America region, 1959
<p>This deposit contains NetCDF files with daily aggregates from Copernicus Era5-Land eight selected indicators, covering the Latin America region, for 1959.</p><p>Each file represents one indicator aggregation for one month of the year. Inside each NetCDF file, the layers contain the daily aggregates.</p><p>For 2m dewpoint pressure, 10m u component of wind, 10m v component of wind, surface pressure, the mean function was used for aggregation. For total precipitation, the sum function was used for aggregation. For 2m temperature, the functions maximum, mean and minimum were used for aggregation.</p><p>Those files were created using the <a href="https://github.com/ErikKusch/KrigR">KrigR</a> package.</p>
ERA5-Land selected indicators daily aggregates for the Latin America region, 1956
<p>This deposit contains NetCDF files with daily aggregates from Copernicus Era5-Land eight selected indicators, covering the Latin America region, for 1956.</p><p>Each file represents one indicator aggregation for one month of the year. Inside each NetCDF file, the layers contain the daily aggregates.</p><p>For 2m dewpoint pressure, 10m u component of wind, 10m v component of wind, surface pressure, the mean function was used for aggregation. For total precipitation, the sum function was used for aggregation. For 2m temperature, the functions maximum, mean and minimum were used for aggregation.</p><p>Those files were created using the <a href="https://github.com/ErikKusch/KrigR">KrigR</a> package.</p>
ERA5-Land selected indicators daily aggregates for the Latin America region, 1957
<p>This deposit contains NetCDF files with daily aggregates from Copernicus Era5-Land eight selected indicators, covering the Latin America region, for 1957.</p><p>Each file represents one indicator aggregation for one month of the year. Inside each NetCDF file, the layers contain the daily aggregates.</p><p>For 2m dewpoint pressure, 10m u component of wind, 10m v component of wind, surface pressure, the mean function was used for aggregation. For total precipitation, the sum function was used for aggregation. For 2m temperature, the functions maximum, mean and minimum were used for aggregation.</p><p>Those files were created using the <a href="https://github.com/ErikKusch/KrigR">KrigR</a> package.</p>
ERA5-Land selected indicators daily aggregates for the Latin America region, 1958
<p>This deposit contains NetCDF files with daily aggregates from Copernicus Era5-Land eight selected indicators, covering the Latin America region, for 1958.</p><p>Each file represents one indicator aggregation for one month of the year. Inside each NetCDF file, the layers contain the daily aggregates.</p><p>For 2m dewpoint pressure, 10m u component of wind, 10m v component of wind, surface pressure, the mean function was used for aggregation. For total precipitation, the sum function was used for aggregation. For 2m temperature, the functions maximum, mean and minimum were used for aggregation.</p><p>Those files were created using the <a href="https://github.com/ErikKusch/KrigR">KrigR</a> package.</p>
ERA5-Land selected indicators daily aggregates for the Latin America region, 1955
<p>This deposit contains NetCDF files with daily aggregates from Copernicus Era5-Land eight selected indicators, covering the Latin America region, for 1955.</p><p>Each file represents one indicator aggregation for one month of the year. Inside each NetCDF file, the layers contain the daily aggregates.</p><p>For 2m dewpoint pressure, 10m u component of wind, 10m v component of wind, surface pressure, the mean function was used for aggregation. For total precipitation, the sum function was used for aggregation. For 2m temperature, the functions maximum, mean and minimum were used for aggregation.</p><p>Those files were created using the <a href="https://github.com/ErikKusch/KrigR">KrigR</a> package.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.