Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,694
datasets available to search
ShareScore release 0.7.1
Dataset results
4,694 results for “Data Analysis”
Analysis of correlation-based biomolecular networks from different omics data by fitting stochastic block models
<p><strong>Baum_et_al_2019_Supplementary_Figures.pdf: </strong>Supplementary Figures S1-S4. Legends are included under each figure.</p> <p><strong>sbm-for-correlation-based-networks-master.zip: </strong>Archived source code of R and Python functions for the analyses and example workflow description at time of publication. Files are maintained at https://gitlab.com/biomodlih/sbm-for-correlation-based-networks and https://gitlab.com/kabaum/sbm-for-correlation-based-networks.</p>
Sample data for analysis of sequence variation in HIV
<p>These are downsampled interleaved paired fastq datasets from Jair et. 2019 (<a href="https://doi.org/10.1371/journal.pone.0214820">https://doi.org/10.1371/journal.pone.0214820</a>). The datasets were prepared by:</p> <ol> <li>Downloading original data from NCBI SRA (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA517147)</li> <li>Trimming contaminating Nextera adapters using trim-galore</li> <li>Mapping reads against nxb2 reference of HIV genome (K03455.1) with BWA MEM</li> <li>Restricting mapped reads to <em>pol</em> gene vicinity (K03455.1:2000-5100)</li> <li>Downsampling mapped data to ~10% of the original with Picard's DownsampleSam</li> <li>Converting BAM to Interleaved Fastq with Picard's SamToFastq</li> <li>Gzipping resultant interleaved paired fastq files</li> </ol>
Data for the publication "Meta-analysis of fecal metagenomes reveals global microbial signatures that are specific for colorectal cancer"
<p>This dataset encompasses all data needed to reproduce the analyses presented in <a href="https://www.nature.com/articles/s41591-019-0406-6">Meta-analysis of fecal metagenomes reveals global microbial signatures that are specific for colorectal cancer</a></p> <p>You can also check the <a href="https://github.com/zellerlab/crc_meta">GitHub repository</a></p>
Data For Scalco et al. Clinicopathological correlates of quantitative Amyloid-B Pathology in the Temporal Cortex: Machine learning analysis of 131 cases from an ADRC
<p>Dataset containing 131 de-identified whole slide images (WSIs) with a respective data dictionary. </p> <p><strong>Paper</strong>: Scalco, R., Oliveira, L.C., Lai, Z. et al. Machine learning quantification of Amyloid-β deposits in the temporal lobe of 131 brain bank cases. acta neuropathol commun 12, 134 (2024). https://doi.org/10.1186/s40478-024-01827-7</p> <p><strong>Details</strong>: A total of 131 .svs. WSIs, de-identified using svs-deidentifier v 0.9.1-beta (https://github.com/pearcetm/svs-deidentifier/releases). Dataset is uploaded in batches due to Zenodo data upload limitations.</p> <p><strong>Slide curation/preparation</strong>: All samples were retrieved from archives of the University of California, Davis Alzheimer’s Disease Center Brain Bank (<a href="https://www.ucdmc.ucdavis.edu/alzheimers/">https://www.ucdmc.ucdavis.edu/alzheimers/</a>). Archival samples analyzed in this study were 5 μm formalin fixed, paraffin embedded sections of the superior and middle temporal gyrus from human brain. The tissue had been previously stained with an amyloid-β antibody (4G8, recognizing residues 17-24, BioLegend, formerly Covance) that were first pretreated with formic acid to rid samples of endogenous protein. All slides were digitized using an Aperio AT2 between 20x and 40x magnification.</p> <p><strong>Code:</strong> Please refer to <a href="https://github.com/ucdrubinet/BrainSec">https://github.com/ucdrubinet/BrainSec</a> and <a href="https://github.com/keiserlab/plaquebox-paper">https://github.com/keiserlab/plaquebox-paper</a></p>
Raw data for analysis of archaeological collections using scoring
<p>The file includes data in the form of attributes identifying the cultural or natural origin of archaeological collections composed of flakes and blades made of flint. The collections attributed to the Lower and Middle Palaeolithic in Poland and Germany. The last collection consists of experimentally produced samples. These collections are held at the Muzeum Śląska Opolskiego in Opole, University of Wrocław, University of Silesia in Katowice, Sosnowiec and Landesmuseum für Vorgeschichte in Halle (Saale). The collected data was used to improve the method hitherto applied to distinguish sets composed of artefacts and pseudo-artefacts using so-called scoring. This work was financially supported by the National Science Centre,<br>Poland (Grant 2020/39/B/HS3/02277).</p>
R scripts for analyzing LiDAR data to assess forest canopy structure and perform Principal Component Analysis (PCA) on derived metrics
<p>This repository contains R scripts for analyzing LiDAR data to assess forest canopy structure and perform Principal Component Analysis (PCA) on spectral and LiDAR-derived metrics. The scripts cover LiDAR data processing, canopy height model (CHM) generation, calculation of forest canopy metrics, and PCA analysis.</p>
Experimental data for fracture toughness analysis of sandstone and granite samples under fluid saturation conditions
<p>This database includes experimental results from mode I fracture toughness (KIC) tests conducted on saturated rock specimens. Three lithologies were studied: a porous siliceous sandstone (Corvio, C) and two high-strength, low-porosity granites (Blanco Mera, BM and Blanco Alba, BA). Tests were conducted at room pressure and temperature using the pseudo-compact tension (pCT) methodology. Seven different fluids were used: deionized water, methanol, NaCl-saturated water, mineral oil, diesel fuel, an acidic HCl solution, and a caustic NaOH solution.</p>
Data accessibility in the chemical sciences: an analysis of recent practice in organic chemistry journals
<div> <p>Data is the analysis of the data outputs of 240 randomly selected research papers from 12 top-ranked journals published in early 2023. We investigate author compliance with recommended (but not compulsory) data policies, whether there is evidence to suggest that authors apply FAIR data guidance in their data publishing, and if the existence of specific recommendations for publishing NMR data by some journals encourages compliance. Files in the data package have been provided in both human and machine-readable forms. The main dataset is available in the Excel file Data worksheet.XLSX, the contents of which can also be found in Main_dataset.CSV, Data_types.CSV, and Article_selection.CSV with explanations of the variable coding used in the studies in Variable_names.CSV, Codes.CSV, and FAIR_variable_coding.CSV. The R code used for the article selection can be found in Article_selection.R. Data about article types from the journals that contain original research data is in Article_types.CSV. Data collected for analysis in our sister paper[4] can be found in Extended_Adherence.CSV, Extended_Crystallography.CSV, Extended_DAS.CSV, Extended_File_Types.CSV, and Extended_Submission_Process.CSV. A full list of files in the data package and a short description for each is given in README.TXT.</p> </div>
Gut Analysis Toolbox: Data and code associated with JCS manuscript
<p>The data and python code in jupyter notebooks are associated with the manuscript: <strong><em>Sorensen et al. Gut Analysis Toolbox: Automating quantitative analysis of enteric neurons. J Cell Sci 2024; jcs.261950. doi: <a href="https://doi.org/10.1242/jcs.261950" target="_blank" rel="noopener">https://doi.org/10.1242/jcs.261950</a></em></strong></p> <ul> <li><strong>FigS1_analysis.zip</strong>: Data files (csv) and jupyter notebooks (ipynb) pertaining to Fig. S1D,E.</li> <li><strong>Fig3_analysis.zip</strong>: Data files (csv) and jupyter notebooks (ipynb) pertaining to Fig. 3D-N. <ul> <li>The images and analysis files associated with analysis in GAT are also uploaded: CalR_CalB_GAT_analysis.zip</li> <li>The images used in this analysis are from EXP174 in this dataset: <a href="https://zenodo.org/records/7236748">https://zenodo.org/records/7236748</a></li> </ul> </li> </ul>
RDF version of the data from Hagar I. Labouta et al. Meta-Analysis of Nanoparticle Cytotoxicity via Data-Mining the Literature. NanoImpact (2019)
<p>This is an RDFied version of the dataset published by Hagar I. Labouta et al. Meta-Analysis of Nanoparticle Cytotoxicity via Data-Mining the Literature. NanoImpact (2019).</p> <p>The original dataset publication DOI: <a href="https://doi.org/10.1021/acsnano.8b07562">https://doi.org/10.1021/acsnano.8b07562</a></p> <p>The Original publication authors: Hagar I. Labouta, Nasimeh Asgarian, Kristina Rinker, and David T. Cramb</p>
Raw data for Infrastructure and Awareness Landscape Analysis in sub-Saharan Africa
<p>Persistent identifiers that are well-connected are essential for enhancing research, researchers, and research institutions. The comprehensive raw data shared on PIDs infrastructure and awareness landscape analysis in sub-Saharan Africa is taken from service providers and organizations, including Open DOAR, the Registry of Open Access Repositories (ROAR), the Registry of Research Data Repositories (Re3data), UNESCO, Lyrasis (Dspace), Dataverse, Open Journal System (OJS), among others. The data shared here was collected in August 2023. The data shared are secondary data, and the position is strictly based on the primary source data author. The data is restricted to what is available on the internet and does not include locally hosted offline data.</p> <p>Further analysis of the raw data suggested some salient implications for PIDs awareness in the region. Variations were observed across the various data sources, while some interesting correlations emerged from the collected data. There are countries with some PIDs infrastructure, while others are yet to establish a visible presence in PIDs infrastructure. The visibility of PIDs is somewhat related to the awareness level as well as the policy established on open access in the represented countries across the region.</p>
Data archive for "Modified rice bran arabinoxylan as a nutraceutical in health and disease — A scoping review with bibliometric analysis"
<p>v1.0.0 Release with the publication of the paper on PLoS One</p> <p>Ooi, S. L., Micalos, P. S., & Pak, S. C. (2023). Modified rice bran arabinoxylan as a nutraceutical in health and disease—A scoping review with bibliometric analysis. PLOS ONE, 18(8), e0290314. https://doi.org/10.1371/journal.pone.0290314</p> <p><strong>Full Changelog</strong>: https://github.com/sooi10/RBACScoping/commits/NetworkAnalysis</p>
Data for paper "Parametric schedulability analysis of a launcher flight control system under reactivity constraints"
<p>This is the data set (models, sources and results) for the paper "Parametric schedulability analysis of a launcher flight control system under reactivity constraints" published in Informatica Fundamentae in 2021.</p>
Van Dijk et al. (2021), A meta-analysis of projected global food demand and population at risk of hunger for the period 2010–2050, data and scripts
<p>This repository contains all data and R scripts to reproduce the figures in Van Dijk et al. (2021), A meta-analysis of global food demand and population at risk of hunger projections for the period 2010-2050, Nature Food. More specifically, it includes two databases: (1) A database with standardized information to describe the characteristics of 57 studies that were identified by the systematic literature review and (2) The Global Food Security Projections Database v1.0.1 with harmonized projections for three global food security indicators: food consumption in kcal per capita and total kcal, and population at risk of hunger. The database also includes projections for total global population that are required to derive the global food security indicators.</p> <p>The two scripts (nf_figures.r and nf_meta_regression.r) can be used to reproduce the figures and tables in the main paper and the supplementary information. Please start with the first script, which sources the second script. </p> <p>This is the first version of the Global Food Projections Database. We expect to update the data, including additional studies and variables in the future. For issues and suggestions, please contact michiel.vandijk@wur.nl.</p> <p> </p>
Data from the parametric analysis of masonry pointed arches with limit analysis subjected to vertical self-weight plus a vertical concentrated live load
<p>For each one of the simulations performed from the parametric analysis of masonry pointed arches with limit analysis, this database contains a .txt, a .vtk and a .png file. In the .txt file the elapsed time and the collapse multiplier of each simulation can be found. The .vtk file contains all the geometry and displacement values of every masonry panel. Finally, the .png file presents the collapse mechanism obtained. </p>
Data from the parametric analysis of masonry pointed arches with limit analysis subjected to vertical self-weight plus a proportional horizontal live load
<p>For each one of the simulations performed from the parametric analysis of masonry pointed arches with limit analysis, this database contains a .txt, a .vtk and a .png file. In the .txt file the elapsed time and the collapse multiplier of each simulation can be found. The .vtk file contains all the geometry and displacement values of every masonry panel. Finally, the .png file presents the collapse mechanism obtained. </p>
Raw and aggregated data for the study introduced in the article "An analysis of citing and referencing habits across all scholarly disciplines: approaches and trends in bibliographic metadata errors"
<p>This dataset contains all the raw data and aggregated data subject of the study introduced in the article "An analysis of citing and referencing habits across all scholarly disciplines: approaches and trends in bibliographic metadata errors". The study is based on the bibliographic and citation data contained in 729 articles published in 147 journals in 27 subject areas. The articles contained a total amount of 34,140 bibliographic references and 55,100 mentions and quotations overall.</p> <p>The dataset is composed of a series of files:</p> <ul> <li>the files "subject_area_<discipline-name>.csv" contain the raw data of the articles published in the journals of all the disciplines considered in the study;</li> <li>the file "article_data_summary.csv" contains the aggregated data created considering the raw data in the previous files, which have been used to creating all the tables and figures in the article;</li> <li>the file "starred_metadata_set.csv" contains information about the most used subset of bibliographic metadata;</li> <li>the file "journals_selection.csv" contains information about all the journals selected for the study.</li> </ul>
Data for Rasch analysis of the three Field Practice Experiences Scales
<p>Data for Rasch analysis of the three Field Practice Experiences Scales. Data rare from Danish teacher education program students. Contains the following variables:</p> <p>Variables O1 to O12 are the items for the Observed scale</p> <p>Variables P1 to P12 are the items for the Practised scale</p> <p>Variables F1 to F12 are the items to the Received feedback scale</p> <p>Gender: 1 = female, 2 = male</p> <p>Campus: 1 = campus A, 2 = campus B</p> <p>T_Progr (teacher education program): 1 = regular, 2 = other</p> <p>P_level (level of latest field practice placement): 1 = 3rd level, 2 = 2nd level, 3 = 1st level</p> <p>Age: 1 = 24 years and younger, 2 = 25 years and older</p>
Data from: Efficacy of labile carbon addition to reduce fast-growing, invasive non-native plants: A review and meta-analysis
<p>Data and analysis in R for the publication "Efficacy of labile carbon addition to reduce fast-growing, invasive non-native plants: A review and meta-analysis" by Ossanna & Gornish (2023), <em>Journal of Applied Ecology</em>, <em>60</em>(2), 218-228. <a href="http://doi.org/10.1111/1365-2664.14324">https://doi.org/10.1111/1365-2664.14324</a>.</p>
Data from: Transcriptomic meta-analysis reveals unannotated long non-coding RNAs related to the immune response in sheep
<p>This dataset contains additional files from the manuscript: "Transcriptomic meta-analysis reveals unannotated long non-coding RNAs related to the immune response in sheep".</p> <p>The files included are:</p> <p>- All novel lncRNA transcript annotation GTF file ( lncrnas.gtf )</p> <p>- High-confidence lncRNA gene annotation GTF file ( lncrnas_evidence.gtf )</p> <p>- All novel lncRNA transcript annotation GTF file remapped to the ARS-UI_Ramb_v2.0 genome ( lncrnas_remapped_v2.gtf )</p> <p>- Raw count estimates of the extended annotation ( rawcounts.csv )</p> <p>- TPM values of the extended annotation ( tpmcounts.csv )</p> <p>- Supplementary data to the published article (.xlsx, .pdf)</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.