Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9,786
datasets available to search
ShareScore release 0.7.1
Dataset results
9,786 results for “selection”
Feature selection on microbial profiles of CRC samples with chopin2 (powered by hdlib)
<p>This Zenodo entry contains the result of the feature selection algorithm implemented through a backward variable elimination strategy in <a href="https://github.com/cumbof/chopin2" target="_blank" rel="noopener">chopin2</a> (powered by <a href="https://github.com/cumbof/hdlib" target="_blank" rel="noopener">hdlib</a>) applied on <a href="https://github.com/biobakery/MetaPhlAn" target="_blank" rel="noopener">MetaPhlAn3</a> microbial profiles of a public dataset of metagenomic stool samples collected from patients affected by the colorectal cancer (CRC) as well as from healthy individuals.</p> <p>Microbial profiles have been extracted through the <a href="https://bioconductor.org/packages/release/data/experiment/html/curatedMetagenomicData.html" target="_blank" rel="noopener">curatedMetagenomicData</a> package for R under the IDs <em>ThomasAM_2018a</em>, <em>ThomasAM_2018b</em>, and <em>ThomasAM_2019_a</em>.</p> <p>The feature selection algorithm is implemented as a backward variable elimination method, and it makes use of the vector-symbolic architecture described in <a href="https://doi.org/10.3390/a13090233" target="_blank" rel="noopener">Cumbo F 2020</a>.</p> <p>Deposited data is described below:</p> <ul> <li><em>datasets.tar.gz</em>: it contains the datasets used as input of <em>chopin2</em> as the result of merging the three datasets with relative abundances mentioned above, also stratified by age and sex (with prefix RA). The same datasets have been also binarized (with prefix BIN);</li> <li><em>hd-models.tar.gz</em>: it contains the output of the feature selection performed with <em>chopin2</em> (powered by <em>hdlib</em>) on the datasets with both relative abundance and binary profiles (RA and BIN);</li> <li><em>ml-models.tar.gz</em>: it contains the result of the feature selection produced with classical wrapper-based techniques (i.e., Random Forest, Decision Tree, Support Vector Machine, Logistic Regression, and Extreme Gradient Boosting) in addition to a Python 3.8 script to reproduce the results.</li> </ul> <p>Please note that the datasets <em>RA__ThomasAM__species.csv</em> and <em>BIN__ThomasAM__species.csv</em> are also included into the <em>datasets.tar.gz</em> archive.</p>
An estimate of fitness reduction from mutation accumulation in a mammal allows assessment of the consequences of relaxed selection: Dataset
<p>Supplementary files (data and analysis) for "An estimate of fitness reduction from mutation accumulation in a mammal allows assessment of the consequences of relaxed selection"</p> <p>Supplementary File 1: C3H_pheno_fix_Jun7_2023_nolowmut.csv</p> <p>Data for all mice in MA experiment including: mouse ID, sire, dam, generation, mating ID, sex, weight at 3 weeks, weight at 6 weeks, tail length, litter size, litter ID, line ID</p> <p> </p> <p>Supplementary File 2: C3H_pheno_Kontrol_June2023.csv</p> <p>Data for all control mice including: mouse ID, sire, dam, generation, mating ID, sex, weight at 3 weeks, weight at 6 weeks, tail length, litter size, litter ID, line ID</p> <p> </p> <p>Supplementary File 3: C3H_birthdates.csv</p> <p>Data for all C3H mice including: mouse ID, birthdate</p> <p> </p> <p>Supplementary File 4: MA_pheno.R</p> <p>R code for visualising trait data, running linear regressions, and comparing control and MA experiment data</p> <p> </p> <p>Supplementary File 5: C3H_pheno_burnin20_Jun7_2023_nolowmut.csv</p> <p>Data for all mice in MA experiment including a 20 generation burn-in to simulate mutation-drift balance for Animal model analyses: mouse ID, sire, dam, generation, mating ID, sex, weight at 3 weeks, weight at 6 weeks, tail length, litter size, litter ID, line ID</p> <p> </p> <p>Supplementary File 6: asreml_C3H_ALL.R</p> <p>R code for estimating mutational heritabilities using mixed model analysis</p> <p> </p> <p>Supplementary File 7: C3H_ped_rekey_Jun2023.csv</p> <p>Pedigree data for all mice in MA experiment</p> <p> </p> <p>Supplementary File 8: C3H_ped_rekey_KEY.csv</p> <p>Key for pedigree data file</p> <p> </p> <p>Supplementary File 9: plot_pedigree_tree_MS_final.R</p> <p>R code for visualising pedigree of mice in MA experiment</p>
Raw and processed hydro-meteorological variables of Jucar river basin for feature selection
<p>The dataset Processed data – input WQEISS.csv was employed for the input variable selection step in Zaniolo et al., 2018. It includes monthly values of 28 hydro-meteorological variables and indexes of Jucar river basin, Spain, for the period 1986-2000, namely:</p> <ul> <li>2 temporal features: day and month of the year;</li> <li>12 inputs to the Jucar State Index: average monthly storage and groundwater levels, average three months river runoff, and cumulated areal precipitation over 12 months;</li> <li>8 additional observed variables in the basin: three months average outflows from, and inflows to, the main reservoirs, and mean monthly areal temperatures;</li> <li>6 traditional drought indicators: Standardized Precipitation Index (SPI) and Standardized Precipitation and Evaporation Index (SPEI). SPI and SPEI indicators are computed on mean monthly data over the entire basin for 3, 6, and 12 months time aggregations.</li> </ul> <p>The last column of the dataset reports the target variable, i.e., the monthly nominal shortage of water conveyed to the irrigation districts simulated via AQUATOOL model. For further details on the dataset please consult Zaniolo et al., 2018, or the dedicated website <a href="http://www.nrm.deib.polimi.it/?page_id=2438">http://www.nrm.deib.polimi.it/?page_id=2438</a></p> <p>The unprocessed data used to compute indices and temporal cumulations in Processed data – input WQEISS.csv are reported in table Raw Data.csv. Public observations of rainfall, streamflows and storage levels come from the SAIH (Hydrological Automatic Information System) of the CHJ (Jucar Hydrological Confederation). Users can directly download data for the last 12 months on the dedicated webpage <a href="http://saih.chj.es/chj/saih/?f">http://saih.chj.es/chj/saih/?f</a> while previous data records are provided for free by CHJ upon request. Observations from piezometers are downloadable from the Piezometric Network Information section section of the CHJ <a href="https://www.chj.es/es-es/medioambiente/redescontrol/Paginas/Piezometr%C3%ADa.aspx">https://www.chj.es/es-es/medioambiente/redescontrol/Paginas/Piezometr%C3%ADa.aspx</a>.</p>
Data to "Object visibility, not energy expenditure, accounts for spatial biases in human grasp selection"
<p>This record contains experimental and analysis scripts (written in Matlab) as well as raw and processed data to reproduce the results shown in:</p> <p><strong>Maiello, G</strong>.<sup> †</sup>, Paulun, V. C.<sup> †</sup>, Klein, L. K. , & Fleming, R. W. (2018) Object visibility, not energy expenditure, accounts for spatial biases in human grasp selection. <em>i-Perception,10</em>(1), 1–5. doi:10.1177/2041669519827608.</p> <p><sup>†</sup>co-first authors</p>
LEUKOS' dataset: A selection of sixteenth and seventeenth century Stambøger from the Royal Danish Library (Copenhagen)
<p>The data were generated to document LEUKOS’ research for objective 1, point 3 (O1.3) of the research project. LEUKOS’ research question (RQ) and brief description of O1: LEUKOS investigates whether and how, during his stay in Wittenberg (1586-1588), Bruno’s notions of the soul and of language were influenced by the new Protestant understandings of philosophy as evidenced by their discussions and suggested reforms of the liberal arts at the philosophical faculty and in Melanchthon’s works, and drew conclusions regarding the relation between Giordano Bruno and the Reformation. To reconstruct the social, intellectual, and political context in Wittenberg during Bruno’s stay (1586-1588) and pave the way for the interpretation of Bruno’s praise of the Reformers as tolerant (O1), LEUKOS reconstructs the environment that Bruno joined in Wittenberg at the time when Aristotle was reintroduced in the curricula of Luther’s university by focusing (O1.1) on the debate concerning the soul in Wittenberg (Gnesio-Lutherans and Philippists) in its philosophical and theological arguments, but also by referring to the political implications of this debate in the view of tolerance; (O2.2) on the role of female theologians and intellectuals in the reformed communities, and in the intellectual world, as the Reformation gave them access to direct participation in cultural and entrepreneurial professions (e.g. as printers); (O1.3) on Danish intellectuals in Bruno’s network as Wittenberg was a centre of diffusion of the Reformation also for the Scandinavian countries. Data generated: pictures of a selection of sixteenth and seventeenth century Stambøger (Alba Amicorum) belonging to the Royal Danish Library collections documenting the transnational intellectual network between Germany and Copenhagen in the second half of the sixteenth century and in the first half of the seventeenth century. The data might be useful to researchers working (1) on the history of the book, (2) on the Royal Danish Library sixteenth and seventeenth century book collections, (3) on Stambøger (Alba Amicorum), (4) on intellectual networks between Germany and Denmark in the sixteenth and seventeenth century.</p> <p>Instrument- or software-specific information needed to interpret the data: It is possible to open a HEIC file on Windows 10 or 11 by: 1. Using the Microsoft Photos App: If your Windows 10 or 11 is up to date, the Microsoft Photos app should already support the HEIC file format. Simply right-click on the HEIC file, select "Open With," and choose "Photos" from the list of apps. The Microsoft Photos app will open the HEIC file and display its contents; 2. Converting HEIC files to JPEG or PNG: If the Microsoft Photos app doesn't support HEIC files on your system, you can convert HEIC files to JPEG or PNG format using an online converter or dedicated software. Search for "HEIC to JPEG converter" or "HEIC to PNG converter" in your preferred search engine to find various online converter tools or downloadable software. Upload your HEIC file to the converter tool or software, select the desired output format (JPEG or PNG), and convert the file. Once converted, you can easily view the resulting JPEG or PNG file using any image viewer or photo app on your Windows 10 or 11 PC; 3. Installing a Third-party HEIC Viewer: If you frequently work with HEIC files, you can also install a dedicated HEIC viewer app from the Microsoft Store or other reputable sources. Search for "HEIC viewer" in the Microsoft Store or other platforms, and look for apps that specifically support viewing HEIC files. Install the chosen HEIC viewer app, and use it to open and view HEIC files on your Windows 10 or 11 PC.</p>
Dataset for simulation studies of fleet vehicle selection in terms of pollutant emissions
<p>The purpose of this dataset is to enable the replication of the research results presented in the article: Szczepański E, Jachimowski R, Rudyk T. Simulation studies of fleet vehicle selection in terms of pollutant emissions. Combustion Engines. 2024;196(1):80-88. https://doi.org/10.19206/CE-169802 - published online: 2023-08-10, which discusses the application of simulation in solving the problem of vehicle selection and determining optimal approaches considering pollutant emissions.</p> <p>Dataset contains:</p> <ul> <li>Readme.txt: description of the dataset</li> <li>InputData.csv: contains the input data used in the model, including data from the COPERT model.</li> <li>OutputOptimization.csv: contains output data</li> <li>OutputSummary.xlsx: contains output data</li> <li>imulation_model_xml.fsx: contains the code of the model in XML format.</li> </ul> <p>The dataset was created as part of the E-Laas project (Energy optimal urban logistics As A Service).<br>Project implemented as part of the call ERA-NET Cofund Urban Accessibility and Connectivity (ENUAC China Call) organized by JPI Urban Europe and the National Natural Science Foundation of China (NSFC). This project has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 875022.<br> E-Laas project is carried out in an international consortium. Project coordinator in Europe: Chalmers University of Technology (Sweden), project coordinator in China: Shanghai University (China), consortium members: Tsinghua University (China), Warsaw University of Technology (Poland), cooperation partners: Stockholms stad, Trafikkontoret (Sweden), ParkUnload (Spain), Metropolis GZM (Poland), Shanghai Urban-Rural Construction and Transportation Department (China), Volvo Group Trucks Technology and Operations (Sweden).<br>- The Chinese part of the project is funded by National Natural Science Foundation of China.<br>- The Swedish part of the project is funded by Swedish Energy Agency.<br>- The Polish part of the project is funded by the National Science Centre, Poland (project no. 2022/04/Y/ST8/00134). The value of the co-financing is PLN 878,107.00. Project duration 27/04/2023 - 26/04/2026 (36 months).</p>
Criteria for prioritizing selection of Mexican maize landrace accessions for conservation in situ or ex situ based on phylogenetic analysis
<p>Data for processed SSR markers in maize accessions. A database in Structured Query Language (SQL) is provided. Please see the text file "READMEmaizeSSR.pdf".</p>
Unintended cation crossover influences CO2 reduction selectivity in Cu-based zero-gap electrolysers
<p>Dataset for the publication "Unintended cation crossover influences CO2 reduction selectivity in Cu-based zero-gap electrolysers"</p>
Beethoven in the House: Selective Encodings of Arrangements of Beethoven's opp. 91, 92, and 93.
<p>This dataset contains selective MEI encodings of a number of arrangements of Beethoven's Opp. 91, 92, and 93. These encodings were prepared in the context of the Beethoven in the House project, jointly funded by AHRC and DFG from 2020 to 2023. It is a slight update on v1.0.0 in better organizing the release assets.</p>
Dataset for publication "Enhancing C≥2 product selectivity in electrochemical CO2 reduction by controlling the microstructure of gas diffusion electrodes"
<p>Data used for publication:</p> <p>Broad topic: electrochemical reduction of CO2 using gas diffusion electrodes and neutral electrolyte</p> <p>Data is devided in subfolders named after the figure of the paper.</p> <p>Raw data, processed data, and Origin/Power Point files are all contained in the subfolders </p> <p>A subfolder corresponding to a sample contains: data from a potentiostat, gas and liquid chromatograms, recording of flow, pressure and temperature, tables of calculated Faradaic efficiency (FE), png image of the FE vs t, zipped raw files.</p> <p>.json file was created using a yadg scheme (https://dgbowl.github.io/yadg/master/index.html), and data was processed by a dgpost scheme (<a href="https://pypi.org/project/dgpost/">https://dgbowl.github.io/dgpost/master/index.html</a>)</p>
An empirical model of the Gaia DR3 selection function
<p>Precomputed maps of the M10 parameter used to predict the completeness of the Gaia DR3 source catalogue.</p> <p><strong>allsky_M10_hpx7.hdf5 </strong>is a tessellation of the whole sky in Galactic nested healpix scheme of order 7.</p> <p><strong>allsky_uniq_10.fits</strong> uses an adaptive resolution and is based on the 'uniq' numbering of healpix tiles. Regions of higher density use a finer resolution, up to order 10, chosen so that every tile contains a minimum of twenty sources to compute M10 from.</p> <p>Paper: https://ui.adsabs.harvard.edu/abs/2023A%26A...669A..55C/abstract</p> <p>Tutorial using the GaiaUnlimited python package to query these maps: https://github.com/gaia-unlimited/gaiaunlimited/blob/main/docs/notebooks/dr3-empirical-completeness.ipynb</p> <p>Documentation for the GaiaUnlimited package: https://gaiaunlimited.readthedocs.io/en/latest/</p>
MATLAB codes for : "Diagnosis and Prognosis of Faults in High-Speed Aeronautical Bearings with a Collaborative Selection Incremental Deep Transfer Learning Approach".
<p>The package contains all the materials needed to reproduce the findings of our paper. The paper is published by MDPI Applied Sciences journal and its details are as follow.</p> <p>Berghout, T.; Benbouzid, M. Diagnosis and Prognosis of Faults in High-Speed Aeronautical Bearings with a Collaborative Selection Incremental Deep Transfer Learning Approach. <em>Appl. Sci.</em> <strong>2023</strong>, <em>13</em>, 10916. https://doi.org/10.3390/app131910916</p> <p>1) Please you need to download the dataset from original link provided by introductory paper (Please read the above paper to find out about the datset used).<br> 2) Put the data in folders "RawData" for both experments.<br> 3) Please run the files for each experiment as provided, in alphabetical order.</p>
Common Raven (Corvus corax) Occupancy Survey and Habitat Selection Data in Cliff Habitat of the Central Appalachian Region, USA, 2009-2010
We identified 24 cliff sites across four states of the Central Appalachian Region of the eastern USA (Kentucky, North Carolina, Virginia, and West Virginia) with known raven occupancy at which to perform occupancy surveys for estimating detection probability and the effects of covariates. We surveyed each cliff site 2-4 times in either 2009 or 2010 and recorded time-to-first detection and time to confirmed cliff occupancy during a two-hour survey. Daily surveys were completed between 06:00 and local solar noon. During each survey, we recorded covariates, including air temperature at survey start time, cloud cover, wind speed, and day of year. We also calculated the distance of the observation point from the cliff being surveyed and the forest cover around the cliff. We also collected data thought to be pertinent for habitat selection by ravens on 26 cliffs occupied by ravens and 26 cliffs deemed unoccupied by ravens in 2010. For each cliff, we measured cliff physiographic characteristics, such as cliff length, cliff height, and occlusion by vegetation, and landscape characteristics, including percent forest and urban cover around the cliff and distances from the cliff to the nearest road and human habitation.
Selection for phenotypic plasticity in Rana sylvatica tadpoles, 1998.
The hypothesis that phenotypic plasticity is an adaptation to environmental variation rests on the two assumptions that plasticity improves the performance of individuals that possess it, and that it evolved in response to selection imposed in heterogeneous environments. The first assumption has been upheld by studies showing the beneficial nature of plasticity. The second assumption is difficult to test since it requires knowing about selection acting in the past. However, it can be tested in its general form by asking whether natural selection currently acts to maintain phenotypic plasticity. We adopted this approach in a study of plastic morphological traits in larvae of the wood frog, Rana sylvatica. First we reared tadpoles in artificial ponds for 18 days, in either the presence or absence of Anax dragonfly larvae (confined within cages to prevent them from killing the tadpoles). These conditioning treatments produced dramatic differences in size and shape: tadpoles from ponds with predators were smaller and had relatively short bodies and deep tail fins. We estimated selection by Anax on the two kinds of tadpoles by testing for non-random mortality in overnight predation trials. Dragonflies imposed strong selection by preferentially killing individuals with relatively shallow and short tail fins, and narrow tail muscles. The same traits that exhibited the strongest plasticity were under the strongest selection, except that tail muscle width exhibited no plasticity but experienced strong increasing selection. A laboratory competition experiment, testing for selection in the absence of predators, showed that tadpoles with deep tail fins grew relatively slowly. In the cattle tanks, where there were also no free predators, the predator-induced phenotype survived more poorly and developed slowly, but this cost was apparently not associated with particular morphological traits. These results indicate that selection is currently promoting morphological plasticity in
CFP01 Fish population on selected watersheds at Konza Prairie (Reformatted to the ecocomDP Design Pattern)
This data package is formatted as an ecocomDP (Ecological Community Data Pattern). For more information on ecocomDP see https://github.com/EDIorg/ecocomDP. This Level 1 data package was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-knz/87/8. The abstract below was extracted from the Level 0 data package and is included for context: Fishes were collected by habitat (pool or riffle) at 6 sites in the Kings Creek watershed with a single-pass electrofishing survey with one person operating the electrofisher and two people dipnetting. Collections were made seasonally.
PVC02 Plant species composition on selected watersheds at Konza Prairie (Reformatted to the ecocomDP Design Pattern)
This data package is formatted as an ecocomDP (Ecological Community Data Pattern). For more information on ecocomDP see https://github.com/EDIorg/ecocomDP. This Level 1 data package was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-knz/69/18. The abstract below was extracted from the Level 0 data package and is included for context: Canopy coverage and frequency were recorded in 20 circular 10 sq m plots. Six treatments were sampled, three ungrazed and three to grazed by native grazers. In each case one of the three watersheds was unburned, another burned annually in April, the third burned every four years in April. In each treatment two soils were sampled: a lower-slope deep fertile nonrocky soil (tully silty clay loam), and a shallow rocky soil (florence cherty silt loam) on level to gently sloping ridges. In 1983 another ungrazed annual burn area '1c' was added 'both tully and florence soils' because original area '1d' appeared aberrant.
SIC01 Isotopic composition of select archived soil cores from Konza Prairie
The concentration and isotopic composition of soil carbon and nitrogen were measured from select archived soil cores originally collected for the NSC01 dataset using an isotope ratio mass spectrometer coupled with an elemental analyzer. These soil cores were collected from the lowlands (25 cm depth) of four experimental watersheds in 1982, 1987, 2002, 2010, and 2015. The four experimental watersheds are 001d, n01b, 020b, and n20b.
PPH01 Phenology of selected plant species at Konza Prairie
Twenty-nine selected species of grasses, forbs, and woody vegetation characteristic of a variety of habitats on Konza Prairie are used for phenological measurements. These species are observed weekly for the entire growing season and changes in their phenological states are recorded. The following phenological states are used for this survey: (1) initiation of growth, (2) first anthesis, (3) duration of anthesis, (4) fruits mature, (5) leaves more than 90% dry.
PVC01 Plant species composition on selected watersheds at Konza Prairie
Canopy coverage and frequency of plant species were estimated visually in 20 circular 10 sq m plots. Six treatments were sampled, three ungrazed and three to be grazed (in the future) by native grazers (bison). In each case, one of the three watersheds was unburned, another burned annually in April, and the third burned every four years in April. In each treatment two soils were sampled: a lower slope deep fertile non-rocky soil (Tully silty clay loam) and a shallow rocky soil (Florence cherty silt loam) on level to gently sloping ridges.
PVC02 Plant species composition on selected watersheds at Konza Prairie
Canopy coverage of all vascular plant species were estimated in 20 circular 10 sq m plots for each of the topographic positions within each included watershed at Konza Prairie.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.