Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,634
datasets available to search
ShareScore release 0.9.0
Dataset results
1,634 results for “Data integration”
Supplementary material 2: World Spider Catalog Bibliographic Data: Publications from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063
Ranked list of journal/publisher exported from the World Spider Catalog 14 October 2014 with total articles by source, cumulative articles, and cucmulative proportion of articles.
Section 5.3 "Task Area 3: Multimodal data linking and integration" Figure 10
<p>Figure 10. Data flow to obtain a multimodal data structure (mmDS) with an overarching graph database (MUGDAT).</p> <p>from NFDI Grant Application, "<strong>National Research Data Infrastructure for Microscopy and Bioimage Analysis</strong>" (NFDI4BIOIMAGE)</p>
S1Data: ChIP-seq Data from Ferrie et. al. "p300 Is an Obligate Integrator of Combinatorial Transcription Factors Inputs"
<p>ChIP data from Ferrie et. al. "p300 Is an Obligate Integrator of Combinatorial Transcription Factors Inputs"</p>
Data: Cutting the costs of coastal protection by integrating vegetation in flood defences.
<p>File: levee_crest_height_reduction_per_country_version_July2021.nc<br>Fields: (1) Crest height reduction m per km along the populated coastline susceptible to flooding (return period = 100 years)<br> (2) Crest height reduction cost saving per country in million USD<sub>2005</sub> PPP along the populated coastline susceptible to flooding (return period = 100 years)<br> (3) Cost savings as percentage of GDP<sub>2005</sub> along the urban populated coastline susceptible to flooding (return period = 100 years)</p> <p>File: transectdata_version_July2021.nc<br> Transectdata of vegetated transects within the study area.<br>Fields: <br>(1) rps = return period <br>(2) fid = id of the transects<br>(3) centroids = coordinates of the transects<br>(4) inun = (1) in area susceptible to flooding<br>(5) urban = (1) in urban area, (0) not in urban area<br>(6) veg_width = derived coastal vegetation belt width along the foreshore<br>(7) veg_type = derived coastal vegetation type along the foreshore (1: salt marshes, 2: mangroves)<br>(8) hsig = Offshore significant wave heights (multiple return periods) corresponding to the transects<br>(9) wave period = Offshore peak wave period (multiple return periods) corresponding to the transects<br>(10) surge = Extreme water level combination of surge and tide (m +MSL) (multiple return periods)<br>(11) veg_z0 = elevation at the start of the vegetated zone (m +MSL)<br>(12) hrms_end_noveg = root mean square wave height at the end of the foreshore (without vegetation) (multiple return periods)<br>(13) hrms_endveg = root mean square wave height at the end of the foreshore (with vegetation) (multiple return periods) <br>(14) pdens_15km = population density derived using buffer of 15 kilometre radius</p>
Data for manuscript: A framework for integrating genomics, microbial traits, and ecosystem biogeochemistry
<p>Support manuscript: A framework for integrating genomics, microbial traits, and ecosystem biogeochemistry. </p> <p>Dataset includes 1) model and analysis, and 2) supplemental data. </p> <p>In the "model_analysis" file, we include the ecosys model source code, the modeling runs, and the modeling results. The detailed introduction is in README.md file. </p> <p><strong>Acknowledgments</strong></p> <p>We thank the EMERGE Biology Integration Institute Coordinators (members listed in Supplementary Information) for project guidance and management. This research is a contribution of the EMERGE Biology Integration Institute, funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070 (V.I.R., R.K.V., S.R.S., M.B.S., E.L.B., and the EMERGE Coordinators). Additional support for individual contributors included the following. Z.L. was additionally supported by Lawrence Livermore National Laboratory under the auspices of the U.S. Department of Energy under contract DE-AC52-07NA27344. W.J.R. was supported by the Belowground Biogeochemistry Scientific Focus Area and U.K. was supported by the Watershed Function Science Area, both funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research under contract no. DE-AC02-05CH11231. G.L.M. was supported by the LLNL "Microbes Persist" Soil Microbiome Scientific Focus Area SCW1632 and an associated KBase award SCW1746. N.J.B. was supported by the US Department of Energy, Office of Science (BER), Early Career Research Program (#FP00005182). B.J.W. was supported by an Australian Research Council Future Fellowship (#FT210100521). J.T. was supported by the Laboratory Directed Research and Development Program of Lawrence Berkeley National Laboratory. </p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council’s grant 4.3-2021-00164. This research used resources of the National Energy Research Scientific Computing Center (NERSC) which is a U.S. Department of Energy Office of Science user facility. This research used the Lawrencium computational cluster resource provided by the IT Division at the Lawrence Berkeley National Laboratory (Supported by the Director, Office of Science, Office of Basic Energy Sciences, of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231). </p> <p><strong>Full list of the EMERGE Biology Integration Institute Coordinators and Affiliations</strong></p> <p>Eoin L. Brodie1,2, Sarah C. Bagby3, Jeffrey P. Chanton4, Jessica G. Ernakovich5, Regis Ferriere6,7, Suzanne B. Hodgkins8, William J. Riley1, Virginia I. Rich8,9, Scott R. Saleska6, Matthew B. Sullivan8,9,10, Ruth K. Varner11, Gene W. Tyson12, Malak M. Tfaily13, Ahmed A. Zayed8,9<br> 1Climate and Ecosystem Sciences Division, Lawrence Berkeley National Laboratory; Berkeley, CA 94720, USA.<br>2Department of Environmental Science, Policy and Management, University of California; Berkeley, CA 94720, USA.<br>3Department of Biology, Case Western Reserve University; Cleveland, OH, USA, 44106<br>4Earth Ocean and Atmospheric Sciences, Florida State University; Tallahassee, FL, USA<br>5Department of Natural Resources and the Environment, University of New Hampshire;<br>Durham, NH, USA 03824<br>6Department of Ecology and Evolutionary Biology, University of Arizona; Tucson, AZ,<br>85721, USA<br>7Institut de Biologie de l’ENS, Université Paris Sciences & Lettres; Paris, 75005, France<br>8Department of Microbiology, The Ohio State University; Columbus, OH, USA, 43210<br>9Center of Microbiome Science, The Ohio State University; Columbus, Ohio 43210, USA.<br>10Department of Civil, Environmental and Geodetic engineering, The Ohio State University; Columbus, Ohio 43210, USA.<br>11Department of Earth Sciences and Institute for the Study of Earth, Oceans and Space, University of New Hampshire; Durham, NH 03824, USA.<br>12Centre for Microbiome Research, School of Biomedical Sciences, Queensland University<br>of Technology (QUT), Translational Research Institute; Woolloongabba, QLD, Australia<br>13Department of Environmental Science, University of Arizona; Tucson, AZ, 85721, USA</p>
SpatialMETA: A Novel Framework for Integrating Spatial Transcriptomics and Metabolomics Data
<p>Multimodal analysis of spatial transcriptomics (ST) and spatial metabolomics (SM) has rapidly advanced for characterizing tissue microenvironments. However, integrating ST and SM data remains challenging due to differing morphologies, resolutions, and batch effects. We developed SpatialMETA (Spatial Metabolomics and Transcriptomics Analysis), a novel method for integrating spatial multi-omics data, which aligns ST and SM to a unified resolution, enables both cross-modal and cross-sample integration to identify ST-SM associated spatial patterns, and provides extensive visualization and analysis capabilities. The datasets for SpatialMETA is avaiable. </p>
Data set for the integrated Climate, Land, Energy and Water systems modelling exercise RCLEWs in OSeMOSYS
<p>This dataset refers to the modelling exercise (version01_210616RCLEWs). The dataset contains the OSeMOSYS code used to run the modelling exercise, the model input data, the scenarios model data files, and the results. The code for the results visualization is available at https://github.com/KTH-dESA/teaching-CLEWs_visualization.</p> <p>This is an update of version 01_210827 available at: https://doi.org/10.5281/zenodo.5293834</p>
Data Set for_Integrating torrefaction of pulp industry sludge with anaerobic digestion to produce biomethane and volatile fatty acids: An example of industrial symbiosis for circular bioeconomy
<p>Industrial symbiosis, which allows the sharing of resources between different industries, could help to improve the overall feasibility of bio-based chemicals production. In that regard, this study focused on integrating the torrefaction of pulp industry sludge with anaerobic digestion. More specifically, anaerobic digestion (AD) of pulp sludge-derived torrefaction condensate (TC) was studied to evaluate the biomethane and volatile fatty acid (VFA) potential. The torrefaction condensate produced at 275 and 300 °C was used in AD. The volatile solid content (VS) was 6.69 and 9.01% for the condensate produced at 275 and 300 °C, respectively. The organic fraction of TC mainly contained acetic acid, 2-furanmethanol, and syringol. The methane yield was in the range of 481–772 mL/g VS for the mesophilic and 401–746 mL/g VS for the thermophilic process, respectively. The VFA yield was in the range of 1.1 to 3.4 g/g VS for mesophilic and from 1.5 to 4.7 g/g VS in thermophilic conditions, when methanogenesis was inhibited. Finally, pulp sludge TC is a feasible feedstock to produce platform chemicals like VFA. However, at higher substrate loading, signs of process inhibition were observed because of the relatively increasing concentration of microbial inhibitors</p>
Dataset linking to the paper "Exploring characteristics of national forest inventories for integration with global space-based forest biomass data"
<p>The dataset links to the study titled “Exploring characteristics of national forest inventories for integration with global space-based forest biomass data”. This study is published in the journal “Science of the Total Environment” and the publication can be found at <a href="https://doi.org/10.1016/j.scitotenv.2022.157788">https://doi.org/10.1016/j.scitotenv.2022.157788</a>. The dataset contains four csv files that were used to produce the results and other figures in the paper. The description of the individual data files contained in the dataset is given below.</p> <p><strong>NFI availability and characteristics data: </strong>The data file “NFI_availability_characteristics.csv” contains data on the total number of NFIs, the NFI extent, and the year of the most recent NFI in countries with NFI as reported in FRA 2020 country reports. The respective data variables in the data file are termed as Number_of_NFI, Latest_NFI_extent_FRA2020, and Latest_NFI_year_FRA2020 (NFI years generally refer to the years of data collection). In addition, the data file contains data on the region and tropical domain per country. The tropical and subtropical countries were considered tropical in the analysis and interpretation of the results. These data were used to produce Figure 2 of the study. ArcMap 10.7.1 was used for this purpose. </p> <p><strong>National biomass intercomparison data: </strong>The data file “national_biomass_intercomparison.csv” contains national forest AGB data for the year 2018 from FRA 2020 and CCI Biomass product that were used in the national biomass intercomparison analysis. The total (tons) and average space-based AGB (tons/ha) are extracted directly from the CCI Biomass Map 2018 for each country included in the study. The processing is done in Python and R environments. The spatial resolution of the map is 100 m. The average FRA AGB data in tons per ha was compiled from FRA 2020 country reports. The total FRA AGB data (tons) was estimated by multiplying each country's average FRA AGB data with FRA forest area data (in ha).</p> <p>The data unit for total AGB was converted from tons to gigaton (Gt) in intercomparison analysis. The total CCI Map AGB estimates used in the analysis are termed as CCI_MAP_AGB_Gt in the data file and the average as CCI_Map_AGB_tons.ha. Similarly, the total FRA AGB data are termed as FRA_AGB_Gt and the average as FRA_AGB_ton.ha. The NFI availability and temporality were also used in intercomparison analysis and this data is termed as Latest_NFI_year_FRA2020 in the data file. The data were used to produce Figure 3 of the study in the R environment.</p> <p><strong>NFI plot design characteristics: </strong>The data file named “NFI_plot_design_characteristics.csv” contains data on variables that were used in the analysis of NFI plot designs in 46 tropical countries. This data file mainly contains the data that was used to produce Figure 4 and Figure 6 in the R environment. The value “uniform” in the sampling_stratification variable means no stratification was used in the sampling design. The variable name “psu” stands for primary sampling unit (both cluster and single plots), “psu_distance_km” for the distance between primary sampling units in km, “cluster_plotdis_m” for the distance between plots in meter in the cluster, “plotsize_ha” for plot (single and cluster plots ) size in ha, “plotshape” for plot shapes (single and cluster plots), “ILUA” for Integrated Land Use Assessment. The data were compiled from the latest NFI design manuals and NFI reports.</p> <p><strong>NFI years: </strong>The data file “NFI_years_tropical_countries_data.csv” contains data on NFI years of the latest NFI in 46 tropical countries that were used to produce Figure 1 using ArcMap 10.7.1. The years generally refer to the last years of data collection. Data were compiled from the latest country NFI design manual or NFI report. This included both ongoing and completed NFI.</p>
Data set for "The annual-hydrogen-yield-climatic-response ratio: evaluating the real-life performance of integrated solar water splitting devices"
<p>This data set was used for the modelling in the article M. Kölbach, O. Höhn, K. Rehfeld, M. Finkbeiner, J. Barry, and M. M. May, “The annual-hydrogen-yield-climatic-response ratio: evaluating the real-life performance of integrated solar water splitting devices”<strong><em>,</em></strong> <em>Sustainable Energy Fuels</em>, <strong>2022</strong>, <strong>6</strong>, 4062-4074, <a href="https://doi.org/10.1039/D2SE00561A">https://doi.org/10.1039/D2SE00561A</a>.</p> <p>It contains the External Quantum Efficiency (EQE) data of a wafer-bonded AlGaAs//Si dual-junction solar cell for several top absorber compositions, angle of incidences, and temperatures modelled using the OPTOS formalism (see <a href="https://doi.org/10.1364/OE.24.0A1083">https://doi.org/10.1364/OE.24.0A1083</a> , <a href="https://doi.org/10.1364/OE.23.0A1720">https://doi.org/10.1364/OE.23.0A1720</a> , and <a href="http://doi.org/10.1109/JPHOTOV.2021.3064562"> https://doi.org/10.1109/JPHOTOV.2021.3064562</a>). Moreover, the data set includes hourly resolved direct and diffuse solar spectra for a location near the Neumayer station in Antarctica (-70.67°/-8.28°) that were modelled using the libRadtran software package for the year 2021 (see <a href="https://doi.org/10.1140/epjconf/e2009-00912-1">https://doi.org/10.1140/epjconf/e2009-00912-1</a> and <a href="http://doi.org/10.5194/acp-5-1855-2005">https://doi.org/10.5194/acp-5-1855-2005</a>). The modelling of the spectra was performed employing the predefined “subarctic summer” and “subarctic winter” atmosphere datasets assuming a tilt angle of 70° and 1-axis tracking. For the sake of simplicity, no cloud cover was assumed over the course of the whole year. Finally, the input files required for modelling the climatic response of solar water splitting devices for the selected location in Antarctica using the “climatic_response_function” of YaSoFo (see <a href="http://doi.org/10.5281/zenodo.5257492">https://doi.org/10.5281/zenodo.5257492</a> for an extended example) are included in the data set.</p>
CMIP6 model vertically-integrated net primary production data
<p>Vertically-integrated net primary production (NPP) data from 12 models that participated in phase six of the Coupled Model Intercomparison Project (CMIP6). All data pulled from the Earth System Grid Federation.</p> <p>All model output was regridded onto a common, regular horizontal grid of 1x1 degrees (360 x 180) in longitude by latitude.</p> <p>Units are mol C per metre squared per second.</p> <p>Models are:</p> <ol> <li>ACCESS-ESM1-5</li> <li>CanESM5</li> <li>CESM2</li> <li>CNRM-ESM2-1</li> <li>GFDL-CM4</li> <li>GFDL-ESM4</li> <li>IPSL-CM6A-LR</li> <li>MIROC-ES2L</li> <li>MPI-ESM1-2-HR</li> <li>MRI-ESM2-0</li> <li>NorESM2</li> <li>UKESM1-0-LL</li> </ol>
Data for Integrated Database, ver.1.0
<p>Analyzed data from MD simulations of small RNA motifs, comparison of predicted and measured NMR observables, performance of common RNA force fields, water models and ion parameters. The data are divided into directories by molecule. Simulations are stored in .dat files while .txt files contain a summary of all. The raw-data folder contains the calculated distances and dihedrals of the defined atoms. The dataset also includes an image to visualize molecule or a pdb file to use it in the Mol* Viewer software.</p>
The data for SKYSURF-5: Probing the Integrated Galaxy Light with a SDSS-SKYSURF Cross-Matched Catalog
<p>The SKYSURF Project (Windhorst et al. 2022) analyzes the extragalactic background light (both directly using sky background measurements and indirectly using galaxy counts) using the HST Archive. While HST images probe faint galaxies unseen by ground-based imaging, its small field of view prevents it from probing the large-scale structure around its observations.</p> <p>To supplement SKYSURF analysis, we cross-match SKYSURF pointings with SDSS observations able to probe the surrounding large-scale environment (Bhatia et al. 2024). The tables in this database include galaxies brighter than r=22.5 AB mag photometrically identified in SDSS, within +/-5 arcmin around a SKYSURF pointing.</p> <p>The tables in the Object_AB directory include information on all SDSS objects, organized by the HST camera and filter of the central pointing.</p> <p>The tables in the IGL directory include the total galaxy counts and integrated galaxy light (down to AB mag=22.5) for all SDSS objects surrounding a given SKYSURF image.</p>
Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research
<p><strong>This is the dataset of the report: Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research</strong></p> <p>It contains the following information of all the papers from ASE, FSE, and ICSE in 2023:</p> <ul> <li>Paper title</li> <li>Keyword</li> <li>Is the source data available and accessible in the paper?</li> <li>If the source data is not available, do the authors explain why?</li> <li>Hosting platforms</li> <li>Access mode</li> <li>License</li> <li>Is their experiment data reused from previous work, or newly generated specifically for this study, or combination of both? </li> <li>Do the authors change/modify their experiment data before experiment?</li> <li>What modifications do they perform?</li> <li>Does the link provide detailed instructions about how to replicate their paper?</li> <li>Does the link contains their complete experiment data, their source code or other materials that are necessary to replicate their experiments?</li> <li>What's the data format inside the link?</li> <li>What's the content of the link?</li> </ul> <p> </p> <p>This is a course project and I collect the data in a rush.</p> <p>If you want to use this dataset and find any error, please contact me ;-)</p> <p>My email: echo.xiangchen@gmail.com</p>
CATCH-EyoU: Exploiting European data and testing the integrated theory of youth active EU citizenship: PIDOP subset reanalysis
<p>This is a subset of the full PIDOP dataset. The derived subset contains cross-sectional survey results from the PIDOP questionnaire survey that were collected in 9 European countries (incl. Turkey) during a period of 16-26 year old in 2011. The data set includes 9060 individual cases. The questionnaire used in the survey is published in Barrett, M. & Zani, B. (Eds.) (2015). <em>Political and civic engagement: Multidisciplinary perspectives.</em> Hove: Routledge (p.519-534).</p>
Data archive for Pepper, Bateson and Nettle, 'Telomeres as integrative markers of exposure to stress and adversity: A systematic review and meta-analysis'
<p>Data archive for the paper 'Telomeres as integrative markers of exposure to stress and adversity: A systematic review and meta-analysis' by Gillian Pepper, Melissa Bateson and Daniel Nettle. This version was uploaded in July 2018 after peer-review in the journal Royal Society Open Science. Compared to earlier version, it incorporates some minor error correction to the dataset, and reflects the revised analyses we performed after peer review. </p> <p>Our protocol and recording guide, which were preregistered on the Open Science Framework in 2016, are also included here, as is our PRISMA diagram.</p> <p>The data file 'unprocessed data' contains the data as extracted from the literature, with associations shown both as provided in the original papers, and converted to correlation coefficients. The algorithms for converting all the different associations to correlation coefficients are described in the flowchart and implemented in the R script 'effect conversion algorithms.r'.</p> <p>The data file 'processed data.csv' is the dataset analysed in the paper. Compared to 'unprocessed data.csv', it excludes: associations from studies of non-human animals; duplicate associations; a small number of associations from studies of medical treatments; and associations considered subparts or subscales of other associations. These exclusions are outlined in Methods section of the paper. In addition, in the processed data file, all correlations are aligned in direction so as to make them comparable (variable 'ValencedEffect'); and all associations are assigned to broad and fine categories.The script 'unprocessed to processed.r' makes the processed data file from the unprocessed one, or you can simply work from the processed one directly. </p> <p>The R script 'telomere metanalysis script RSOS REVISED.r' reproduces the analyses found in the paper.</p> <p>This version of the archive (July 17 2018) contains one small correction in the data files compared to all earlier versions. </p>
Supplementary material to the manuscript: Regionalised Heat Demand and Power-To-Heat Capacities in Germany - An Open Data Set for Assessing Renewable Energy Integration
<p>This is the supplementary material for the manuscript:</p> <p>"Regionalised Heat Demand and Power-To-Heat Capacities in Germany - an Open Data Set for Assessing Renewable Energy Integration"</p> <p>Article DOI: <a href="https://doi.org/10.1016/j.apenergy.2019.114161">https://doi.org/10.1016/j.apenergy.2019.114161</a></p> <p>Open access preprint: <a href="https://arxiv.org/abs/1912.03763">https://arxiv.org/abs/1912.03763</a></p> <p> </p> <p><strong>DESCRIPTION OF THE DATASET AND LICENSES:</strong></p> <p>The subdirectory "04_results" contains the regionalised heat demand an power-to-heat capacity data on administrative district level (NUTS-3) for Germany. The subdirectories "01_census_special_evaluation_data" and "02_other_input_data" contain the utilised input data. The subdirectory "03_code" contains the developed and applied source code.</p> <p>The data in this repository are provided under open source licenses. For license information and other general information on the supplementary material, refer to the LICENSE files and README files in the respective subdirectories.</p> <p>For a detailed description of the approach developed by the author, the input data used and the generated results, refer to the manuscript "Regionalised Heat Demand and Power-To-Heat Capacities in Germany - an Open Data Set for Assessing Renewable Energy Integration".</p> <p><strong>METADATA:</strong></p> <p>Sector: Residential Buildings – Space Heating and Domestic Hot Water</p> <p>Geographical scope: Germany</p> <p>Geographical resolution: Administrative districts (NUTS-3)</p> <p>Temporal scope: 2011, three scenarios for 2030</p> <p>Temporal resolution: 15min</p> <p> </p> <p><strong>UNITS:</strong></p> <p>In the final results folders (04_results/01_installed_heating_p2h_capacity; 04_results/02_daily_time_series; 04_results/03_yearly_time_series) the units of the data are indicated in the file names or the column names, e.g. by "in_MW". In case of unit indication in the file name, the unit refers to all columns in the file.</p> <p>In the intermediate results folder (04_results/00_sql_tables_exported_to_csv) all units referring to power are "kW" and all units referring to energy are "kWh".</p> <p><strong>NEWS AND CONTACT:</strong></p> <p>This dataset will be used as part of the <a href="https://wiki.openmod-initiative.org/wiki/Region4FLEX">region4FLEX model</a>. We are currently enhancing the data by temporally and spatially resolved COP time series and determining load shifting potentials. If you wish to receive news or have general questions please contact: wilko.heitkoetter@dlr.de. </p>
Integration of data sets from different sources for modeling gender violence and perception of insecurity
<p>The dataset is composed of three distinct files which aggregate processed data derived from open datasets of three cities: Dublin, San Francisco, and Valencia. The data has been mapped to a grid of 25m² for Valencia and 50m² for Dublin and San Francisco. The respective files are named DATA_ES_VLC.csv, DATA_IE_DUB.csv, and DATA_US_SFO.csv. Additionally, there is a dataset for tweets named DATA_TWT.csv, which contains tweets collected through web scraping and analysed using natural language processing (NLP) algorithms and neural networks. The aim is to identify and classify tweets that discuss gender-based violence in the city of Valencia. Another file, MAP_ES_VLC.csv, includes points collected during various mapathons conducted by the Polytechnic University of Valencia campus for a science project aimed at identifying potentially insecure locations.</p>
Integration of expression datasets to identify biomarkers for accurate Gleason scoring in Prostate Cancer -- Supplementary data
<p>This dataset contains expression data from multiple sources used to identify biomarker candidates for prostate cancer aggressiveness. The data includes transcriptional expression levels, patient metadata, and other relevant features utilized in our machine-learning models. The training dataset was extracted from the repository described at Matos-Filipe, et al. (2022) [1].</p> <p>Raw ML metrics from models gaussian Naïve Bayes classifiers using this dataset are available in </p> <p> </p> <p>[1] <span><span><span>Matos-Filipe, P, et al. "</span></span></span>The usage of transcriptomics datasets as sources of Real-World Data for clinical trialling". <span>bioRxiv (</span><span>2022). </span><span><span>doi:</span> https://doi.org/10.1101/2022.11.10.515995</span></p>
RDF version of the data from Choi, JS. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources (2018)
<p>This is an RDFied version of the dataset published in Choi, JS., Ha, M.K., Trinh, T.X. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources. Sci Rep 8, 6110 (2018)</p> <p>The original dataset publication DOI: <a href="https://doi.org/10.1038/s41598-018-24483-z">https://doi.org/10.1038/s41598-018-24483-z</a></p> <p>The Original publication authors: Jang-Sik Choi, My Kieu Ha, Tung Xuan Trinh, Tae Hyun Yoon & Hyung-Gi Byun</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.