Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
599
datasets available to search
ShareScore release 0.9.0
Dataset results
599 results for “health data”
Sample based prevalence data complementing the European Union One Health 2021 Zoonoses Report - Luxembourg
<p>This dataset contains monitoring data on zoonoses and zoonotic agents under the Directive 2003/99/EC. This Directive requires Member Sates (MSs) to collect, evaluate and report data on zoonoses and zoonotic agents. MSs can also report monitoring data and information on some other pathogenic microbiological agents in foodstuffs. Relevant EU legislation: Commission Regulation (EC) No 2073/2005,Commission Regulation (EC) No 1441/2007, Commission Regulation (EU) No 1086/2011, Commission Regulation (EU) No 209/2013, Commission Regulation(EU) No 217/2014.</p>
Animal disease data complementing the European Union One Health 2021 Zoonoses Report
<p>This dataset contains the mandatory annual data reported for bovine tuberculosis and for bovine and ovine and caprine brucellosis based on Directive 2003/99.</p>
Sample based prevalence data complementing the European Union One Health 2021 Zoonoses Report - Ireland
<p>This dataset contains monitoring data on zoonoses and zoonotic agents under the Directive 2003/99/EC. This Directive requires Member Sates (MSs) to collect, evaluate and report data on zoonoses and zoonotic agents. MSs can also report monitoring data and information on some other pathogenic microbiological agents in foodstuffs. Relevant EU legislation: Commission Regulation (EC) No 2073/2005,Commission Regulation (EC) No 1441/2007, Commission Regulation (EU) No 1086/2011, Commission Regulation (EU) No 209/2013, Commission Regulation(EU) No 217/2014.</p>
The Influence of Land Surface Temperature on Mental Health in Lausanne - Data
<p>Dataset contains LST data for Lausanne used for the report.</p> <p>GeoPackage is in ESPG:21781 and provides coordinated in ESPG:4326 as well.</p> <p>Data contains mean and median LST for Summer (June - August) 1998, 2008 and 2018 and the delta. The delta is calculated by subtractig the data from 2018 (for example deltalstmedian0818 = lst_2018 - lst_2008).</p> <p>Find the code here: https://github.com/aamir-s18/InfluenceLSTGAF</p>
Data and Code for "A planetary health innovation for disease, food, and water challenges in Africa"
<p>Data and R code for both economic and non-economic analysis for the paper titled, "A planetary health innovation for disease, food, and water challenges in Africa"</p>
Data for "Structure and drivers of social networks and their links with health in older adults"
<p>Data, supplementary material and scripts for the paper "Structure and drivers of social networks and their links with health in older adults"</p> <p>Social network is an important factor in promoting healthy aging. However, the mechanisms linking social capital to health are complex. Moreover, most of the social network analysis studies on older adults consider only participants’ relationships and not how these relationships are themselves connected. In this study, we went further than current ego-centered network studies by determining global social network metrics and the structure of relationships among older adult participants of the RECORD Cohort using the Veritas-Social questionnaire. The aim of this study is to identify key dimensions of social networks of older adults, and to evaluate how these dimensions relate to depressive symptoms, life satisfaction, and well-being. Using Principal Component Analyses (PCA), we identified four social network dimensions with psychological meanings. Dimension 1 (homophily) was positively linked with perceived accessibility to services in one’s residential neighborhood but negatively linked with the level of study. Dimension 2 (social integration) as Dimension 3 (social support) was only linked to the number of people living with ego. Dimension 4 was linked with perceived accessibility to local services. Finally, and rather surprisingly, we found that none of the four network dimensions, even the degree, was linked to the three health status metrics.</p>
Raw_data_Combining several indicators to assess the effectiveness of tailor-made health plans in pig farms
<p>Raw data for a paper submitted to Peer Community In Animal Science. </p>
Data for: Evaluation of antibody kinetics and durability in health individuals vaccinated with inactivated COVID-19 vaccine (CoronaVac): a cross-sectional and cohort study in Zhejiang, China
<p><strong>Background</strong>: Although inactivated COVID-19 vaccines are proven to be safe and effective in the general population, the dynamic response and duration of antibodies after vaccination in the real world should be further assessed.</p> <p><strong>Methods</strong>: We enrolled 1067 volunteers who had been vaccinated with one or two doses of CoronaVac in Zhejiang Province, China. Another 90 healthy adults without previous vaccinations were recruited and vaccinated with three doses of CoronaVac, 28 days and 6 months apart. Serum samples were collected from multiple timepoints and analyzed for specific IgM/IgG and neutralizing antibodies (NAbs) for immunogenicity evaluation. Antibody responses to the Delta and Omicron variants were measured by pseudovirus-based neutralization tests.</p> <p><strong>Results</strong>: Our results revealed that binding antibody IgM peaked 14–28 days after one dose of CoronaVac, while IgG and NAbs peaked approximately 1 month after the second dose and then declined slightly over time. Antibody responses had waned by month 6 after vaccination and became undetectable in the majority of individuals at 12 months. Levels of NAbs to live SARS-CoV-2 were correlated with anti-SARS-CoV-2 IgG and NAbs to pseudovirus, but not IgM. Homologous booster around 6 months after primary vaccination activated anamnestic immunity and raised NAbs 25.5-fold. The neutralized fraction subsequently rose to 36.0% for Delta (p=0.03) and 4.3% for Omicron (p=0.004), and the response rate for Omicron rose from 7.9% (7/89) to 17.8% (16/90).</p> <p><strong>Conclusions</strong>: Two doses of CoronaVac vaccine resulted in limited protection over a short duration. The inactivated vaccine booster can reverse the decrease of antibody levels to prime strain, but it does not elicit potent neutralization against Omicron; therefore, the optimization of booster procedures is vital.</p>
Data from: Evaluation of a community health worker home visit intervention to improve child development in South Africa: A cluster-randomized controlled trial
<p><span>This dataset was collected as part of a</span><span> cluster-randomized controlled trial that evaluated the impact of </span>a home visit intervention on child development<span> in Limpopo Province, South Africa. </span><span>Household survey data were collected at baseline and endline. In a subsample of children, neural function was assessed at a lab at endline and at two interim time points. Primary outcomes were: height-for-age z-scores (HAZ) and stunting; child development scores measured using the Malawi Developmental Assessment Tool (MDAT); absolute electroencephalography (EEG) gamma and total power; relative EEG gamma power; and saccadic reaction time (SRT)—</span>an eye-tracking measure of visual processing speed.</p>
Cancer Health Disparities drivers with BERTopic Modelling and PyCaret Evaluation - Text data
<p>The complex interplay of social, behavioral, lifestyle, environmental, health system, and natural health variables contribute to disparities in cancer treatment across racial and ethnic groups. Consequently, it is necessary to identify the variables contributing to cancer health inequalities and develop strategies to achieve health equality. PubMed abstract on Cancer health disparities was scraped with a bio.Entrez python package. Preprocessed data with regex and Natural tool kit (NLTK), topic modelling with BERTopic embeddings, and c-TF-IDF to construct dense clusters and analyze top topics linked with Cancer health disparities. Model evaluation with PyCaret coherence score and web app deployment with Streamlit. The results showed that Topic 32 with terms obese, female, male, school, survey, student, poet, and discrepancy had the best coherence score of 0.3687. In contrast, topic 8, with terms prevalence, adult, income, high, usage, diabetes, education, elderly, change and low, received the least coherence score of 0.3255. The model classifies each Subject Word score based on the scores, the granular topic concerns and trends related to cancer health disparities, investigates the connection between drivers of cancer health disparities, and evaluates the model with their coherence score values</p>
Figure 2 in First data on water mite (Acari, Hydrachnidia) assemblages of Point Rosa Marsh, Harrison Township, Michigan, USA, and their use as environmental bioindicators of aquatic health
Figure 2 Point Rosa Marsh along Lake St. Clair, Harrison Township, Michigan, USA. (A) Taken at
Figure 1 in First data on water mite (Acari, Hydrachnidia) assemblages of Point Rosa Marsh, Harrison Township, Michigan, USA, and their use as environmental bioindicators of aquatic health
Figure 1 Map of Lake St. Clair Metropark with inset showing placement in the Lake St. Clair
Unlocking the Potential of Health Data: A Distributed Analysis Approach based on Personal Health Train Infrastructure
<p>While there is a great availability of medical datasets, they are usually focused on a specific research question. This is useful for making experiments transparent and reproducible, however, these datasets can be more efficiently used in other kinds of analyses, where it not for data privacy issues. The Personal Health Train (PHT) provides a distributed analysis infrastructure that follows the FAIR principles and gives control to the data owners (providers) about how their data are used by scientists or other users (consumers).</p>
Geo-gender-based analysis of human health data
<p>Data supporting analysis in "Geo-gender-based analysis of human health: the presence of cut flower farms can attenuate pesticide exposure in communities, yet women remain most vulnerable"</p>
Data from: Changes in environment and management practices improve foot health in zoo-housed flamingos
<p><strong>Summary</strong></p> <p>This dataset accompanies the publication <strong>"Changes in Environment and Management Practices Improve Foot Health in Zoo-Housed Flamingos"</strong> published in <em>Animals</em>. This study tracked changes in foot lesions for an individual flock of Chilean flamingos (97 birds) at Dublin Zoo (Ireland) over an 18-month period in response to management and substrate changes .</p> <p>Photos of each flamingo's feet were taken on May 6th 2021, when all flamingos had access to their outdoor habitat (<strong>Time Point A</strong>). Photos were taken again on 16th April 2022, following a six month period when the flamingos were restricted to their indoor habitat due to a Government order to prevent the spread of Avian Influenza (<strong>Time Point B</strong>). Final photos were taken on 9th November 2022, six months following the release of the birds back into their outdoor habitat (<strong>Time Point C</strong>). Further details can be found in the corresponding publication. </p> <p>Scoring was undertaken blindly by two independent and trained evaluators. These scores reflect the scoring metric developed by Nielsen et al. 2010, and include the four types of common flamingo foot lesion: hyperkeratosis, fissures, nodular lesions, and papillomatous growths. The independently calculated foot scores were subsequently compared, and in instances where the foot scores did not match, a consensus was sought between both evaluators to provide a final value for subsequent analysis. The data presented here reflects the consensus values used in the analysis. Discrepancies in the foot scores between both evaluators are reported and discussed in the corresponding publication. </p> <p><br> <strong>Description of the Dataset</strong></p> <p>One file is provided in .csv format. The file contains the following 11 columns: </p> <ul> <li><strong>Time_Point:</strong> The Time Point at which photos were taken (A = 6th May 2021, B = 16th April 2022, and C = 9th November 2022). </li> <li><strong>Animal_Identifier: </strong>An anonymous code used to identify individual flamingos (n = 97).</li> <li><strong>Hyperkeratosis_Total: </strong>The total hyperkeratosis score for that flamingo at that Time Point (considering both feet).</li> <li><strong>Fissures_Total: </strong>The total fissures score for that flamingo at that Time Point (considering both feet).</li> <li><strong>Nodular_Lesions_Total: </strong>The total nodular lesions score for that flamingo at that Time Point (considering both feet).</li> <li><strong>Papillomatous_Growths_Total:</strong> The total papillomatous growths score for that flamingo at that Time Point (considering both feet).</li> <li><strong>L_Total:</strong> The total left foot lesions score for that flamingo at that Time Point (considering all types of foot lesion).</li> <li><strong>R_Total:</strong> The total right foot lesions score for that flamingo at that Time Point (considering all types of foot lesion).</li> <li><strong>Overall_Total:</strong> The total foot lesions score for that flamingo at that Time Point (considering both feet and all types of foot lesion).</li> <li><strong>Sex:</strong> The sex of the flamingo (Male or Female) </li> <li><strong>Age: </strong>The age of the flamingo (Years)</li> </ul> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We acknowledge and thank all Dublin Zoo staff and volunteers for their support and assistance throughout the project. Additionally, we thank Dr. Laura Kane for her technical assistance and support. </p> <p> </p> <p><strong>Disclaimer</strong></p> <p>Despite our best efforts at screening the data for errors and inconsistencies, some information could be erroneous. </p> <p> </p> <p><strong>Credit</strong></p> <p>If you use this dataset, please cite the corresponding publication:</p> <p>Mooney, A., McCall, K., Bastow, S., & Rose, P. (2023). Changes in Environment and Management Practices Improve Foot Health in Zoo-Housed Flamingos. <em>Animals, 13</em>(15),<em> </em>2483. <a href="https://doi.org/10.3390/ani13152483">https://doi.org/10.3390/ani13152483</a></p>
Data analysis for "Wellbeing, loneliness, health-related quality of life and perception of technology of older adults in Slovenian senior homes"
<p>We present datasets data analysis conducted in R for the article"Wellbeing, loneliness, health-related quality of life and perception of technology of older adults in Slovenian senior homes".</p>
Cell Health Data Supplementary Files
<p>This dataset contains supplementary files related to the <em><a href="https://github.com/WayScience/cell-health-data">cell-health-data</a></em> repository that were too large to upload on GitHub. Each of these files were saved to and loaded from an external hard drive during Cell Health Data Processing. These files include:</p> <ul> <li>single_cell_classification_probabilities/ : Single cell classification probabilities derived with each model trained in <a href="https://github.com/WayScience/phenotypic_profiling_model">phenotypic_profiling_model</a>. Each compressed csv file in single_cell_classification_probabilities/ corresponds to single-cell classifications from a particular plate as derived with a particular model. The model used to derive the probabilities is indicated by the file's parent folders. Each compressed csv file includes the metadata (location and perturbation) and phenotypic class probabilities for each cell in the plate.</li> <li>classification_profiles/ : This folder contains aggregated data known as "classification profiles". These profiles are generated from the single-cell classification probabilities found in the 'single_cell_classification_probabilities/' folder. Specifically, for each model, we find the mean of the single-cell classification probabilities across each perturbation and cell line to create a composite profile. This aggregated data provides a summarized view of cell behavior for each perturbation/cell line combination, as predicted by each model.</li> </ul> <p>For more information regarding the generation of this data, please see <a href="https://github.com/WayScience/cell-health-data/tree/master/4.classify-single-cell-phenotypes">cell-health-data/4.classify-single-cell-phenotypes</a>.</p>
Data for: Detection of oomycete pathogens in UK peat-free growing media and implications for plant health
<p>This dataset on Zenodo accompanies the manuscript Frederickson-Matika <em>et al.</em> (2024), Detection of oomycete pathogens in UK peat-free growing media and implications for plant health.</p> <p>There are two files:</p> <ul> <li>metadata.tsv - plain text table as tab-separated variables</li> <li>raw_data.tar.gz - compressed archive of 43 paired raw FASTQ files</li> </ul> <p>This represents a subset of two complete Illumina MiSeq plates (in two dated folderes) run at the James Hutton Institute containing other environmental samples using the same protocol. Only the synthetic controls and peat-free samples are provided here.<br><br>To repeat the analysis described in the paper, first install THAPBI PICT. See <a href="https://github.com/peterjc/thapbi-pict/">https://github.com/peterjc/thapbi-pict/ </a>for instructions. At the time of the paper, v1.0.14 was the current release.</p> <p>Next, decompress the raw data into a folder of paired gzipped FASTQ files. There is no need to decompress those:</p> <pre><code> $ tar -zxvf raw_data.tar.gz<br> $ ls -1 plate_20220505/ plate_20230608/</code></pre> <p>If you wish, verify the checksums to confirm the data integrity:</p> <pre><code> $ cd plate_20220505/ $ md5sum -c MD5SUM.txt<br> $ cd ../plate_20230608/ $ md5sum -c MD5SUM.txt<br> $ cd ..</code></pre> <p>Setup output directories:</p> <pre><code><code> $ mkdir -p intermediate/ summary/</code></code></pre> <pre>Run the THAPBI PICT pipeline:</pre> <pre><code> $ thapbi_pict pipeline -m 1s3g \<br> -i plate_*/ -o summary/peat-free \<br> -y plate_*/GBL*.fastq.gz \<br> -n plate_*/GBL*.fastq.gz \<br> -s intermediate/ \<br> -t metadata.tsv -u \<br> -x 9 -c 1,2,3,4,5,6,7,8</code><br><br></pre> <p>The options here are as follows:</p> <ul> <li>-i - two input directories of paired raw FASTQ files.</li> <li>-n - negative controls used to increase the absolute abundance threshold</li> <li>-y - synthetic controls used to increase the fractional abundance threshold</li> <li>-s - optional location to store intermediate files</li> <li>-o - output stem for reports</li> <li>-t - filename for tab-separated-variable metadata</li> <li>-u - show unsequenced samples defined in the metadata</li> <li>-x - which metadata column contains Illumina FASTQ filename stems</li> <li>-c - which metadata columns to include in the report.</li> </ul> <p>This assumes the following key default settings:</p> <ul> <li>-a 100 (default absolite abundance threshold)</li> <li>-f 0.001 (default fractional abundance threshold)</li> <li>-d -(default provided ITS1 database).</li> </ul> <p>With these settings, only synthetic sequences were found in the controls, and therefore the thresholds were not automatically increased any further.</p> <p>Opening the output file summary/peat-free.ITS1.samples.1s3g.xlsx in Excel or similar should show you a table resembling Table 1 in the paper, but one row per sequencing sample, and additional columns with per-sample per-species read counts etc.</p>
Feasibility of Monitoring Health Data in Pediatric Patients Undergoing Chemotherapy
ClinicalTrials.gov study NCT04134429. IPD Sharing: YES. Countries: 1. Publications: 1.
Assessing the Performance of Artificial Intelligence (AI)-Augmented Electronic Health Record (EHR) Data Abstraction for Clinical Trial Patient Screening
ClinicalTrials.gov study NCT06561217. IPD Sharing: NO. Countries: 1. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.