Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,940
datasets available to search
ShareScore release 0.9.0
Dataset results
1,940 results for “data sample”
Data for manuscript, "An optimized workflow for MS-based quantitative proteomics of challenging clinical bronchoalveolar lavage fluid (BALF) samples"
<p>Clinical BALF samples are rich in biomolecules, including proteins, and useful for molecular studies of lung health and disease. However, MS based proteomic analysis of BALF is impeded by the dynamic range of protein abundance, and potential for interfering contaminants. We have developed a workflow that eliminates these challenges. By combining high abundance protein depletion, protein trapping, clean-up, and in-situ tryptic digestion, our workflow is compatible with both qualitative and quantitative MS-based proteomic analysis. The workflow includes collection of endogenous peptides for peptidomic analysis of BALF, if desired, as well as amenability to offline semi-preparative or microscale fractionation of peptide mixtures prior to LC-MS/MS analysis, for increased depth of analysis. We show the effectiveness of this workflow on BALF samples from COPD patients. Overall, our workflow should allow MS-based proteomics to be applied to a wide variety of studies focused on BALF clinical samples. </p> <p>Note: Due to the nature of some of the files, file <em>wendt005_ostr0103_18260_20210831_BALF_FAIMS_MS2_TMT16.msf, wendt005_ostr0103_18976_20230202_quantReport.msf, cmsptc_higgi022_18988_20230203_18976DW_EnF_hcdlT_1R.raw, cmsptc_higgi022_18988_20230203_18976DW_EnF_hcdlT_2R.raw, cmsptc_higgi022_18988_20230203_18976DW_EnF_hcdlT_3R.raw and cmsptc_higgi022_18988_20230203_18976DW_Eclipse_noFAIMS_quantReport.msf</em> were zipped into compressed folders before uploading.</p>
Data of Nutrition Knowledge in a Sample of College Students in Jordan
<p>This cross-sectional research assessed nutrition knowledge in a sample of college students in Amman, Jordan and its association with food security and other risk factors. An Arabic Nutrition Knowledge Index (ANKI) was developed and validated in 122 college students (study-1). In study-2, the demographics scale, validated ANKI, and Arabic Individual Food Insecurity Experience Scale (FIES) were administered to 470 students from the same university. </p> <p><strong>Ethical considerations:</strong>This manuscript has been read and approved by all authors. The authors confirm that there are no other persons, who satisfied the criteria for authorship, but are not listed. The order of authors listed in the manuscript has been approved by all of them. They also understand that the Corresponding Author is the sole contact for the Editorial process, and holds the responsibility for communicating with the other author about progress, submissions of revisions and final approval of proofs. Moreover, the authors declare that this manuscript is original, has not been published before, and is not currently being considered for publication elsewhere. Furthermore, all data used in the study is confidential and the lead author has full access to the data reported in the manuscript. We confirm that there are no known conflicts of interest associated with this publication, which did not receive any financial support. Finally, the reporting of this work is compliant with The Code of Ethics of the World Medical Association (Declaration of Helsinki). In addition to that, the protocol of this research is approved by the Institutional Review Board at the University of Jordan, Amman, Jordan (ref no.: 2021-89).</p> <p> </p>
Raw data drosophila parasitoids sampled in 2008 life history traits
<p>Sampling of parasitoids in 2008 and 2011</p> <p>Life history and physiological traits measurements at different temperatures : developement time, fecundity, mass, metabolic rate, etc.</p> <p> </p>
Data from: Quantifying and estimating ecological network diversity based on incomplete sampling data
<p>An ecological network refers to the ecological interactions among sets of species. Quantification of ecological network diversity and related sampling/estimation challenges have explicit analogues in species diversity research. A unified framework based on Hill numbers and their generalizations was developed to quantify taxonomic, phylogenetic, and functional diversity. Drawing on this unified framework, we propose three dimensions of network diversity that incorporate the frequency (or strength) of interactions, species' phylogenies and traits. As with surveys in species inventories, nearly all network studies are based on sampling data and thus also suffer from under-sampling effects. Adapting the sampling/estimation theory and the iNEXT (interpolation/extrapolation) standardization developed for species diversity research, we propose the iNEXT.link method to analyze network sampling data. The proposed method integrates the following four inference procedures: (i) Assessment of sample completeness of networks, (ii) asymptotic analysis via estimating the true network diversity, (iii) non-asymptotic analysis based on standardizing sample completeness via rarefaction and extrapolation with network diversity, and (iv) estimation of the degree of unevenness or specialization in networks based on standardized diversity. Interaction data between European trees and saproxylic beetles are used for illustrating the proposed procedures. The software iNEXT.link is developed to facilitate all computations and graphics.</p>
Data sample INDIP Validation Paper
<p>Standardized data from Mobilise-D HYA dataset (one participant, data from both laboratory and 2.5 hours) and Mobilise-D technical validation dataset (one subject for each cohort - HOA, PD, MS, COPD, CHF, PFF; data from 2.5 hours) are provided in the shared folder, as an example of the the work proposed in the publication "A multi-sensor wearable system for the assessment of diseased gait in real-world conditions" that has been accepted for publication in Frontiers Journal. Please refer to that publication for further information. Please cite that publication if using these data. </p>
Supplementary data for: Tackling hysteresis in conformational sampling --- how to be forgetful with MEMENTO
<p>This is the supplementary data for our publication entitled: "Tackling hysteresis in conformational sampling --- how to be forgetful with MEMENTO". This version corresponds to the first revision of our paper in response to reviewer feedback.</p> <p>Abstract:</p> <p><<<</p> <p>The structure of proteins has long been recognised to hold the key to understanding and engineering their function, and rapid advances in structural biology and protein structure prediction are now supplying researchers with an ever-increasing wealth of structural information. Most of the time, however, structures can only be determined in free energy minima, one at a time. While conformational flexibility may thus be inferred from static end-state structures, their interconversion mechanisms --- a central ambition of structural biology --- are often beyond the scope of direct experimentation. Given the dynamical nature of the processes in question, many studies have attempted to explore conformational transitions using molecular dynamics (MD). However, ensuring proper convergence and reversibility in the predicted transitions is extremely challenging. In particular, a commonly used technique to map out a path from a starting to a target conformation called steered MD (SMD) can suffer from starting-state dependence (hysteresis) when combined with techniques such as umbrella sampling (US) to compute the free energy profile of a transition. \\</p> <p> </p> <p>Here, we study this problem in detail on conformational changes of increasing complexity. We also present a new, history-independent approach that we term “MEMENTO” (Morphing End states by Modelling Ensembles with iNdependent TOpologies) to generate paths that alleviate hysteresis in the construction of conformational free energy profiles. MEMENTO utilises template-based structure modelling to restore physically reasonable protein conformations based on coordinate interpolation (morphing) as an ensemble of plausible intermediates, from which a smooth path is picked. We compare SMD and MEMENTO on well-characterized test cases (the toy peptide deca-alanine and the enzyme adenylate kinase) before discussing its use in more complicated systems (the kinase P38α and the bacterial leucine transporter LeuT). Our work shows that for all but the simplest systems SMD paths should not in general be used to seed umbrella sampling or related techniques, unless the paths are validated by consistent results from biased runs in opposite directions. MEMENTO, on the other hand performs well as a flexible tool to generate intermediate structures for umbrella sampling. We also demonstrate that extended end-state sampling combined with MEMENTO can aid the discovery of collective variables on a case-by-case basis.</p> <p>>>></p> <p>There are 4 folders in this dataset for the 4 simulation systems studied in the paper. Deca-alanine, ADK, P38a and LEUT. Each folder contains key coordinate files used to initialise MEMENTO and SMD runs, and data produced in the course of simulations (as labelled in the folders). In the simulation folders, 'raw_data' corresponds to the outputs produced by PLUMED during REUS or SMD, 'convergence' contains WHAM outputs taken at a certain percentage of the data included, and 'pointsfile' are the metadata files required by the Grossfield WHAM implementation.</p>
Training data for the "Computational textural mapping harmonises sampling variation and reveals multidimensional histopathological fingerprints"
<p>There are two ZIP-files consisting of small histological image tiles that have been used to detect and quantify distinct tissue textures and lymphocyte proportions from H&E-stained clear cell renal cell carcinoma (KIRC) digital tissue sections of the Cancer Genome Atlas (TCGA) image archive and the Helsinki dataset.</p> <p>The <strong>tissue_classification </strong>file contains 300x300px tissue texture image tiles (n=52,713) representing renal cancer (“cancer”; n=13,057, 24.8%); normal renal (“normal”; n=8,652, 16.4%); stromal (“stroma”; n= 5,460, 10.4%) including smooth muscle, fibrous stroma and blood vessels; red blood cells (“blood”; n=996, 1.9%); empty background (“empty”; n=16,026, 30.4%); and other textures including necrotic, torn and adipose tissue (“other”; n=8,522, 16.2%). Image tiles have been randomly selected from the TCGA-KIRC WSI and the Helsinki datasets.</p> <p>The <strong>binary_lymphocytes </strong>file contains mostly 256x256px-sized but also smaller image tiles of Low (n=20,092, 80.1%) or High (n=5,003, 19.9%) lymphocyte density (n=25,095). Image tiles have been randomly selected from the TCGA-KIRC WSI dataset.</p> <p>All accuracy of all annotations have been double-checked. However, the classification between multiple tissue textures or lymphocyte density can be sometimes ambiguous.</p> <p>The deep learning model parameters trained with the ResNet-18 infrastructure for (1) lymphocyte and (2) texture classification are named as (1) <strong>resnet18_binary_lymphocytes.pth</strong> and (2) <strong>resnet18_tissue_classification.pth</strong>. Codes and instructions to use these are found in <a href="https://github.com/vahvero/RCC_textures_and_lymphocytes_publication_image_analysis">https://github.com/vahvero/RCC_textures_and_lymphocytes_publication_image_analysis</a>.</p> <p> </p> <p>If you use either work, please cite the publication by Brummer O et al (1) AND the TCGA Research Network (2):<br><strong>(1) </strong><strong>Brummer, O., Pölönen, P., Mustjoki, S. <em>et al.</em> Computational textural mapping harmonises sampling variation and reveals multidimensional histopathological fingerprints. <em>Br J Cancer</em> 129, 683–695 (2023). </strong><a href="https://doi.org/10.1038/s41416-023-02329-4">https://doi.org/10.1038/s41416-023-02329-4</a></p> <p><strong>(2) The results shown here are in whole or part based upon data generated by the TCGA Research Network: </strong><strong><a href="https://www.cancer.gov/tcga">https://www.cancer.gov/tcga</a></strong><strong>.</strong></p>
GSD Sample Data
<p>GSD Open Source Sample Data as found in the Beeline software package that was likely generated using a Boolean Network method described in Ríos, O., Frias, S., Rodríguez, A. <em>et al.</em> (2015). All gene names have been normalized to their NCBI identifiers and any NCBI IDs that contain parentheses such as NCBI:12796(-KTS) or NCBI:12796(+KTS) denoted a sequence variant described by the string within the parentheses.</p>
MultiScent-20 Data - Binary responses (correct vs. incorrect answers) data from full sample, for each odour
<p>Data used for IRT analysis in the article</p> <p>"The Multiscent-20: A Digital Odour Identification Test Developed with Item<br> Response Theory", with the following authors:</p> <p>Marcio Nakanishi 1, *, Pedro Renato de Paula Brandão 2, *, Gustavo Subtil Magalhães<br> Freire 1 , Luis Gustavo do Amaral Vinha 3 , Marco Aurélio Fornazieri 4 , Wilma Terezinha<br> Anselmo Lima 5 , Claudia Galvão 6 , Thomas Hummel 7</p> <p> </p> <p>1 Department of Otorhinolaryngology, University Hospital of Brasília, University of<br> Brasília, Brasília, DF, Brazil<br> 2 Neurology Department, Hospital Sírio-Libanês, Brasília, DF, Brazil<br> 3 Statistics Department, University of Brasília, DF, Brazil<br> 4 State University of Londrina, Pontifical Catholic University of Paraná<br> 5 Department of Ophthalmology and Otorhinolaryngology, Ribeirão Preto Medical<br> School-University of São Paulo, Brazil<br> 6 NOAR Brasil Ltda, São Paulo-SP, Brazil<br> 7 Smell and Taste Clinic, Department of Otorhinolaryngology, Technical University of<br> Dresden, Dresden, Germany</p> <p> </p> <p> </p> <p>MultiScent-20 descrition:</p> <p>The Multiscent-20 is a portable tablet computer with an integrated hardware<br> system consisting of a central processing unit, touchscreen, Wi-Fi antenna, USB dock,<br> power supply, and rechargeable battery. Additionally, it contains an odour system<br> composed of 20 microcartridges, an air filter, a system that generates a dry air stream,<br> and an odour-dispensing opening. The device is capable of presenting 20 different<br> odours from individual odour capsules, which store the olfactory stimuli by incorporating<br> the odours into an oil-resistant polymer. The capsules are loaded through an insertion<br> port on the back of the device.<br> A software-controlled processing unit governs the flow of dry air, generating a<br> constant air stream that passes through the capsule, releasing individual odours<br> through a small opening at the upper front of the device. In this study, each scent was<br> presented for 5 seconds, with an interval of at least 6 seconds between odour<br> presentations. Each capsule contained 35 µL of oil-based odour solution, allowing the<br> device to maintain consistent odour intensity and identifiability for up to 100 activations,<br> as per the manufacturer's instructions. The Givaudan Flavours and Fragrance<br> Corporation (São Paulo, Brazil) prepared the odours used in the capsules.<br> A digital application (software) was developed to present odours and record<br> olfactory function test results. The device delivers odours through a dry air system,<br> leaving no residue in the environment or on users. Users can access the test using an<br> application on the device (screen version) or mobile phone. The results are accessible<br> through the integrated software, allowing for storage and analysis of the data.</p>
Data from the manuscript 'Accurate detection of shared genetic architecture from GWAS summary statistics in the small-sample context'
<p>Data sets from the manuscript 'Accurate detection of shared genetic architecture from GWAS summary statistics in the small-sample context'. These include the test statistics from analyses of real and simulated data, and the data used to generate the figures relating to the goodness-of-fit of the generalised extreme value distribution to the GPS test statistics under the null. Please see the enclosed README for more details.</p>
HoloFood Data Portal Samples
<p>HoloFood data portal table with full list of samples contained within the portal</p>
(EMPIR 19ENG06 HEFMAG) Data sets of measurements of magnetic loss and complex permeability on amorphous and nanocrystalline samples up to the MHz range
<p>We measured the magnetic losses and the complex permeability of amorphous and nanocrystalline ribbons from DC to 1 GHz by combined application of fluxmetric and transmission line methods. Two transverse field annealed Co-based amorphous alloys, ~13 µm and ~25 µm tick and two nanocrystalline Finemet type alloys, ~13 µm and ~20 µm tick, endowed with defined transverse magnetic anisotropy, were characterized. </p>
Data for figures in "Next-generation ice nucleating particle sampling on aircraft: Characterization of the High-volume flow aERosol particle filter sAmpler (HERA)"
<p>Atmospheric ice nucleating particle (INP) concentration data from the free troposphere are sparse, but urgently needed to understand vertical transport processes of INPs and their influence on cloud formation and properties. Here, we introduce the new High-volume flow aERosol particle filter sAmpler (HERA) which was specially developed for installation on research aircraft and subsequent offline INP analysis. HERA is a modular system constisting of a sampling unit and a powerful pump unit and has several features which were integrated specifically for INP sampling. Firstly, the pump unit enables sampling at flow rates exceeding 100 L min<sup>−1</sup>, which is well above typical flow rates of aircraft INP sampling systems described in the literature (~10 L min<sup>−1</sup>). Consequently, required sampling times to capture rare, high-temperature INPs (≥-15 °C) are reduced in comparison to other systems and potential source regions of INPs can be confined more precisely. Secondly, the sampling unit is designed as a seven-way valve, enabling switching between six filter holders and a bypass with one filter being sampled at a time. In contrast to other aircraft INP sampling systems, the valve position is controlled remotely via software so that manual filter changes in-flight are eliminated and the potential for sample contamination is decreased. This design is compatible with a high degree of automation, i.e., triggering filter changes depending on parameters like flight altitude, geographical location, temperature, or time. In addition to the design and principle of operation of HERA, this paper presents laboratory characterization experiments with size-selected test substances, i.e., SNOMAX® and Arizona Test Dust. The particles were sampled on filters with HERA, varying either particle diameter (300 nm to 800 nm) or flow rate (10 L min<sup>−1</sup> to 100 L min<sup>−1</sup>) between experiments. The subsequent offline INP analysis showed good agreement with literature data and comparable sampling efficiencies for all investigated particle sizes and flow rates. Furthermore, the deposition efficiency of atmospheric INPs in HERA was compared to a straightforward filter sampler and good agreement was found. Finally, results from the first campaign of HERA on the High Altitude and LOng range research aircraft (HALO) demonstrate the functionality of the new system in the context of aircraft application.</p> <p>The given csv files contain the data for reproducing the figures in the publication. The data structure of the csv files is explained in the README file.</p>
Credit scoring with class imbalance data: An out-of-sample and out-of-time perspective
<p>The raw datasets provided here are intended for use in a Data in Brief article. These comprehensive files, sourced from the Freddie Mac website, offer quarterly snapshots of mortgage loans that have been originated in the USA since 1999, along with details of their subsequent repayment behaviours. This data remains current and is updated every three months. Specifically, the loan origination data present here encompasses amortized fixed-rate mortgage loans from 1999 up to June 2022. In contrast, the performance data is presented on a monthly basis, detailing loan repayment profiles from 1999 until September 30, 2022. Both the origination and performance datasets feature a unique loan ID, which can be utilized to integrate the data on loan originations with that of loan repayments.</p>
Sample data for analysis of demographic potential of the 15-minute city in northern and southern France
<pre>This upload contains two Geopackage files of raw data used for urban analysis in the outskirts of Lille and Nice, France. <br>The data include building footprints (layer "building"), roads (layer "road"), and administrative boundaries (layer "adm_boundaries")<br>extracted from version 3.3 of the French dataset BD TOPO®3 (IGN, 2023) for the municipalities of Santes, Hallennes-lez-Haubourdin,<br>Haubourdin, and Emmerin in northern France (Geopackage "DPC_59.gpkg") and Drap, Cantaron and La Trinité in southern France <br>(Geopackage "DPC_06.gpkg").</pre> <pre> </pre> <pre>Metadata for these layers is available here: https://geoservices.ign.fr/sites/default/files/2023-01/DC_BDTOPO_3-3.pdf</pre> <pre> </pre> <pre>Additionally, this upload contains the results of the following algorithms available in GitHub (<a href="https://github.com/perezjoan/emc2-WP2?tab=readme-ov-file">https://github.com/perezjoan/emc2-WP2?tab=readme-ov-file</a>)</pre> <pre><code> </code></pre> <pre>1. The<code> </code>identification<code> </code>of<code> </code>main<code> </code>streets using the QGIS plugin Morpheo (layers "road_morpheo" and "buffer_morpheo") <br><a href="https://plugins.qgis.org/plugins/morpheo/">https://plugins.qgis.org/plugins/morpheo/</a> </pre> <pre><code>2. </code>The<code> </code>identification of main streets in local contexts – connectivity locally weighted<code> </code>(layer "road_LocRelCon")</pre> <pre><code>3. </code>Basic morphometry<code> </code>of<code> </code>buildings<code> </code>(layer "building_morpho")</pre> <pre><code>4. </code>Evaluation<code> </code>of<code> </code>the<code> </code>number<code> </code>of<code> </code>dwellings<code> </code>within<code> </code>inhabited<code> </code>buildings<code> </code>(layer "building_dwellings")</pre> <pre>5. Projecting<code> </code>population<code> </code>potential<code> </code>accessible from<code> </code>main<code> </code>streets<code> </code>(layer "road_pop_results")</pre> <pre> </pre> <pre>Project website: <a href="http://emc2-dut.org/">http://emc2-dut.org/</a></pre> <pre> </pre> <pre>Publications using this sample data: <br>Perez, J. and Fusco, G., 2024. Potential of the 15-Minute Peripheral City: Identifying Main Streets and Population Within Walking Distance. In: O. Gervasi, B. Murgante, C. Garau, D. Taniar, A.M.A.C. Rocha and M.N. Faginas Lago, eds. <em>Computational Science and Its Applications – ICCSA 2024 Workshops. ICCSA 2024</em>. Lecture Notes in Computer Science, vol 14817. Cham: Springer, pp.50-60. <a href="https://doi.org/10.1007/978-3-031-65238-7_4" target="_blank" rel="noopener noreferrer">https://doi.org/10.1007/978-3-031-65238-7_4</a>.</pre> <p><strong>Acknowledgement.</strong> <a name="_Hlk162443883"></a>This work is part of the emc2 project, which received the grant ANR-23-DUTP-0003-01 from the French National Research Agency (ANR) within the DUT Partnership.</p>
Data from: Reduced sampling intensity through key sampling site selection for optimal characterization of riverine fish communities by eDNA metabarcoding
Open the record for dataset details and reuse information.
Data from: Flying high: Sampling savanna vegetation with UAV-lidar
Open the record for dataset details and reuse information.
Code for: A century of wild bee sampling: historical data and neural network analysis reveal ecological traits associated with species loss
Open the record for dataset details and reuse information.
Data for: PickMe: Sample selection for species tree reconstruction using coalescent weighted quartets
Open the record for dataset details and reuse information.
Data from: Skyline fossilized birth-death model is robust to violations of sampling assumptions in total-evidence dating
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.