Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11,174
datasets available to search
ShareScore release 0.7.1
Dataset results
11,174 results for “identifiers”
Mycorrhizal fungal communities identified from seedlings planted in the Taylor, Dalton, and Boundary fire complexes which burned in 2004
This dataset contains the operational taxonomic unit table and taxonomic assignments for fungi that were associated with the roots of seedlings planted into the 2004 burn sites. There were 458 seedlings from 22 of the 32 established intensive sites (Johnstone and Hollingsworth 2019) consistenting of black spruce, white spruce, aspen, and lodgepole pine.
Dataset of "Electronic structure and defect states in bismuth and antimony sulphides identified by energy-resolved electrochemical impedance spectroscopy"
Understanding the nature of the defects in the absorber materials, namely point defects, their formation mechanism and the contribution to the properties is essential for the photovoltaic device performance improvement. They are one the reasons why chalcogenide-based solar cells do not yet meet expected high power conversion efficiencies. Here we identify and present energy distribution of defects in Bi2S3 and Sb2S3, and their (SbxBi(100-x))2S3 alloys (with x = 0, 10, 33, 50, 67, 90, 100 at% Sb content) chalcogenides, being explored for emerging photovoltaic applications as they are earth-abundant and highly absorbing in the visible light range. We show that their density of states (DOS) and related parameters can be obtained experimentally by energy-resolved electrochemical impedance spectroscopy (ER-EIS) in a technically simple and quick way, where ER-EIS data are well correlated with theoretical DFT calculations. ER-EIS reveals that in Bi2S3 there are only shallow defects at CBM. In Sb2S3, ER-EIS reveals also midgap states which can be the cause of low electrical conductivity of Sb2S3. We also explain the discrepancy in the reported values of ionisation potentials and the bandgaps of the Bi- and Sb-chalcogenides. Dominant sulphur vacancy defect was identified in Bi- and Sb-chalcogenides whereas in ternary (SbxBi(100-x))2S3 system, merely 10 at.% of Bi transforms the midgap sulphur defects to shallow ones. This provides novel strategy for healing the midgap defects in Sb2S3, which is crucial for boosting the PV performance and tuning the electrical conductivity in Sb2S3.
Geochemical Characterizations for Identifying Fugitive Dust Deposition and Enrichment of Surface and Subsurface Subalpine Soils from Phosphorus Mining, Eastern Ashley National Forest, Utah, 2022-2023.
Phosphorus is a non-renewable resource essential for all life. Anthropogenic alterations to the phosphorus cycle have led to widespread phosphorus pollution, and the unsustainable management of P has led to the threat of global depletion of phosphorus resources. Thus, accounting for the natural and anthropogenic flow paths of phosphorus is essential for its conservation and pollution reduction. One such source of human alteration to the phosphorus-cycle is phosphate rock mining. Mining, however, has many adverse environmental effects, including widespread fugitive dust emissions. Dust collection in the Ashley National Forest of northeastern Utah, proximate to a surface phosphorus mine, has shown phosphorus concentrations in dust more than four times that of other regional samples. Elevated phosphorus in dust near active surface mining suggests that mining emissions may alter the natural phosphorus loading of the soils in the National Forest through dust deposition; however, no research has been done to identify the abundance and range of mine-attributable phosphorus enrichment in the soils surrounding phosphate mining activities. The combined geospatial and geochemical approach of this study shows that surface soil phosphorus concentrations were found to be enriched above naturally occurring levels up to 6.5 km from mining activity (enrichment factor > 1.5), with the most significant enrichment occurring within the first 3 km (enrichment factor > 2). On average, surface phosphorus concentrations were significantly enriched by 25% within 6.5 km of phosphorus mining activity. Observed phosphorus enrichment was positively correlated with the presence of fluorapatite in the soil, which is the primary phosphorus-mineral extracted from the nearby mine. Further, bioavailable phosphorus concentrations were also higher for the soils that were enriched in phosphorus. This study shows that fugitive emissions associated with the surface mining of phosphate rock are a significan
Salvaging the Internet Hate Machine: Using the discourse of extremist online subcultures to identify emergent extreme speech
<p>This dataset accompanies a paper submitted to the WebSci 20 conference. In this paper, we present a lexicon of 'extreme speech' that may be used to detect hate speech and extreme speech on online platforms. We outline a cross-disciplinary research protocol through which this lexicon is initially extracted from a corpus of 3,335,265 posts from 4chan's /pol/ sub-forum using a hybrid method comprising word2vec modeling and subsequent snowballing of nearest neighbours of a small initial expert seed list of extreme language. The choice of corpus is significant, as 4chan is a space of rapid language innovation and obscure extreme vernacular, complicating generalised approaches. Our lexicon detects significantly more extreme posts within a corpus from a more mainstream platform (Reddit) than another popular lexicon, Hatebase, with similar accuracy. Our lexicon and the method of its creation thus provide a contribution to the study of the toxicity of online subcultures similar to 4chan, as well as more mainstream platforms. As we demonstrate, the lexicon allows for more effective detecting of extreme speech in these spaces. This method and the lexicon have further been made available through an open-source web tool for the study of online social platforms, 4CAT. The computational methods and lexicon on offer here can thus be used by a wide academic audience, fostering interdisciplinary approaches to the study of online hate and extreme speech. </p> <p>The dataset comprises the following items:</p> <ul> <li>The 4chan corpus from which the extreme speech lexicon was generated (posts from /pol/, 1 October 2019 - 1 November 2019)</li> <li>The Reddit corpus used to verify and test the lexicon (posts from the_donald, theredpill, politics and chapotraphouse, 1 October 2019 - 1 November 2019)</li> <li>The word2vec model from which the extreme speech lexicon was generated</li> <li>The extreme speech lexicon that was generated</li> </ul>
PopDel identifies medium-size deletions jointly in tens of thousands of genomes - Variant call sets
<p>This data set contains the variant calls sets generated by different tools for the benchmarks in the paper <a href="https://www.nature.com/articles/s41467-020-20850-5">PopDel identifies medium-size deletions simultaneously in tens of thousands of genomes</a>. It includes the VCFs/BCFs for the following test cases:</p> <ul> <li>Random deletion simulation on up to 1000 chromosome 21 samples</li> <li>1000 Genomes Project deletions inserted into simulated chromosomes 17 to 22 of up to 500 samples</li> <li>HG001 (NA12878)</li> <li>Trio of <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG002_NA24385_son/NIST_HiSeq_HG002_Homogeneity-10953946/">HG002</a> + <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG003_NA24149_father/NIST_HiSeq_HG003_Homogeneity-12389378/">HG003</a> + <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG004_NA24143_mother/NIST_HiSeq_HG004_Homogeneity-14572558/">HG004</a></li> <li><a href="https://github.com/Illumina/Polaris/wiki/HiSeqX-Diversity-Cohort">Polaris Diversity cohort</a></li> <li><a href="https://github.com/Illumina/Polaris/wiki/HiSeqX-Kids-Cohort">Polaris Kids cohort</a></li> </ul> <p>Further, the long and short read reference call sets for HG001 are provided. For HG002 the reference call set and the high confidence regions by the Genome in a Bottle consortium are provided.</p> <p>For details on how the files have been created, please refer to the paper and the script repository on <a href="https://github.com/kehrlab/PopDel-scripts">GitHub</a>.</p>
Bulk and single-cell gene expression profiling of SARS-CoV-2 infected human cell lines identifies molecular targets for therapeutic intervention
<p>Single cell RNA seq datasets used for analysis in the Bulk and single-cell gene expression profiling of SARS-CoV-2 infected human cell lines identifies molecular targets for therapeutic intervention</p>
S100 | PFASREACH | List of PFAS identified in REACH 2019
<p>This is the collection associated with list S100 PFASREACH List of PFAS identified in REACH 2019 on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>A list of 437 PFAS identified in Registration, Evaluation, Authorisation and Restriction of Chemicals <a href="https://eur-lex.europa.eu/eli/reg/2008/1272">(REACH) Reg. (EC) No 1272/2008</a> in September 2019. Of these, 17 are produced at >1000 tonnes/year and 84 at >10 tonnes per year for use in Europe. Collaborative effort between Hans Peter Arp (NGI) and Emma Schymanski (LCSB) within <a href="https://zenodo.org/communities/zeropm-h2020?page=1&size=20">ZeroPM</a> (EU H2020 grant 101036756).</p> <p>Updates: 14 Dec 2022: added 5 new CIDs following PubChem deposition. 16 April 2025: removed CID <a href="https://pubchem.ncbi.nlm.nih.gov/compound/10996402">10996402</a> and <a href="https://pubchem.ncbi.nlm.nih.gov/compound/21967041">21967041</a> as they are not PFAS, due to external report via PubChem. </p>
Saharan HPEs identified and described in Armon et al. (in rev.)
<p>This dataset represents heavy precipitation events (HPEs) identified and described in Armon et al. (in rev., Weather and Climate Extremes). </p><p>It contains two files, both similar to Table 1 in the manuscript: </p><p>(a) 'Sahara_HPEs.csv' contains details of 41,978 HPEs throughout the Sahara</p><p>(b) 'Sahara_Core_HPEs.csv' contains details of 650 HPEs identified over the core of the Sahara.</p><p>Please note, coordinates represent the centre-of-mass of precipitation for every HPE.</p>
Potential forest conservation value rasters for Denmark from Assmann et al. "LiDAR data fusion and machine learning identify temperate forests of high conservation value"
<p>Potential forest conservation value (high / low) rasters for Denmark based on a remote sensing data fusion approach. Please see manuscript (below) for a detailed description of the methods and data products. </p> <p><br>Jakob J. Assmann, Pil B. M. Pedersen, Jesper E. Moeslund, Cornelius Senf, Urs A. Treier, Derek Corcoran, Zsófia Koma, Thomas Nord-Larsen, Signe Normand. In prep. LiDAR data fusion and machine learning identify temperate forests of high conservation value.</p> <p><br>When using the data, please cite the above manuscript. </p> <p><br>Files description:</p> <ul> <li>Compressed and cloud optimised rasters of potential forest conservation value projections for Denmark (10 m res.) in EPSG:3857 <ul> <li>forest_quality_ranger_biowide_10m_cog_epsg3857.tif RandomForest model projections based on BIOWIDE stratification (!! best performing model !!)</li> <li>forest_quality_ranger_sustainscapes_10m_cog_epsg3857.tif RandomForest model projections based on SustainScapes stratification</li> <li>forest_quality_gbm_biowide_10m_cog_epsg3857.tif GBM model projections based on BIOWIDE stratification</li> <li>forest_quality_gbm_sustainscapes_10m_cog_epsg3857.tif GBM model projections based on SustainScapes stratification</li> </ul> </li> </ul> <p> </p> <ul> <li>Aggregated rasters of potential forest conservation value projections for Denmark (100 m res.) in EPSG:25832 <ul> <li>forest_quality_ranger_biowide_100m.tif RandomForest model projections based on BIOWIDE stratification (!! best performing model !!)</li> <li>forest_quality_ranger_sustainscapes_100m.tif RandomForest model projections based on SustainScapes stratification</li> <li>forest_quality_gbm_biowide_100m.tif GBM model projections based on BIOWIDE stratification</li> <li>forest_quality_gbm_sustainscapes_100m.tif GBM model projections based on SustainScapes stratification </li> </ul> </li> </ul> <p> </p> <ul> <li>Uncompressed and tiled rasters of potential forest conservation value projections for Denmark (10 m res.) in EPSG:25832<br>Please note: the archives contain approx. 42k tiles, each 10 x 10 km, as well as a VRT file for covenient loading. <ul> <li>forest_quality_ranger_biowide_10m.zip RandomForest model projections based on BIOWIDE stratification (!! best performing model !!)</li> <li>forest_quality_ranger_sustainscapes_10m.zip RandomForest model projections based on SustainScapes stratification</li> <li>forest_quality_gbm_biowide_10m.zip GBM model projections based on BIOWIDE stratification</li> <li>forest_quality_gbm_sustainscapes_10m.zip GBM model projections based on SustainScapes stratification</li> </ul> </li> </ul>
OpenCitations Meta RDF dataset of identifiers metadata and its provenance information
<p>This dataset is a specialized subset of the OpenCitations Meta RDF data, focusing exclusively on data related to <strong>identifiers </strong>(<a href="http://purl.org/spar/datacite/Identifier" target="_blank" rel="noopener">http://purl.org/spar/datacite/Identifier</a>) of bibliographic resources. It contains all the metadata and its provenance information, structured specifically around identifiers, in JSON-LD format.</p> <p>The inner folders are named through the <strong>supplier prefix</strong> of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to <strong>06*0</strong>).</p> <p>After that, the folders have <strong>numeric names</strong>, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the <strong>zipped </strong>RDF data.</p> <p>At the same level, additional folders containing the <strong>provenance </strong>are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called <strong>prov</strong>, also in zipped JSON-LD format.</p> <p>For example, data related to the entity is located in the folder /id/06250/10000/1000/1000.zip, while information about provenance in /id/06250/10000/1000/prov/se.zip</p> <p>Additional information about OpenCitations Meta at the <a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p>
Identifying the mechanisms by which irrigation can cool urban green spaces in summer
<p>This dataset contains the measured soil moisture and microclimate data from two (2021 and 2022) urban green space irrigation experiments conducted in Burnley, Melbourne, Australia. The experiments consisted of two treatments, irrigated turf and unirrigated turf. The purpose of the experiments was to provide testing (2021) and evaluation (2022) data for an urban ecohydrological model, UT&C. </p> <p><br>After evaluating the performance of UT&C in modelling soil moisture and microclimate, UT&C was used to model the surface energy balance and evapotranspiration processes of the irrigated and unirrigated turf. This dataset also contains the modelled soil moisture, microclimate, surface energy balance and evapotranspiration data, as well as the measured background climate data at the reference climate station and the forcing data for the model.</p> <p><br>The aims of this study were to:<br>i) identify the proportional contribution of different evapotranspiration processes to irrigation cooling effect, and <br>ii) quantify the impacts of different irrigation amounts (from 2 to 30 mm/d) on the cooling effect of irrigating turfgrass in Melbourne, Australia during normal summer conditions.</p> <p>This study was published in:<br>Pui Kwan Cheung, Naika Meili, Kerry A. Nice, Stephen J. Livesley (2024). Identifying the mechanisms by which irrigation can cool urban green spaces in summer. Urban Climate. 55,101914. https://doi.org/10.1016/j.uclim.2024.101914.</p>
Metabomatching: Using Genetic Association to Identify Metabolites in Proton NMR Spectroscopy. CoLaus Pseudospectra.
<p>Summary statistics between urine NMR metabolome features and genotypes in the CoLaus cohort. Used as test pseudospectra for metabomatching, a method for metabolite identification using genetic spiking.</p>
Metabomatching: Using Genetic Association to Identify Metabolites in Proton NMR Spectroscopy. SHIP Pseudospectra.
<p>Summary statistics between urine NMR metabolome features and genotypes in the SHIP cohort. Used as test pseudospectra for metabomatching, a method for metabolite identification using genetic spiking.</p>
Dataset to "Persistent Identifiers for File Formats: enabling preservation and re-use of research data"
<p>This fileset includes a "preprint" and the main dataset <em>fileformatRecognizer</em> (as .xlsx and .csv) to the paper "Persistent identifiers for file formats: enabling preservation and re-use of research data" submitted to iPRES 2019, but subsequently rejected after peer review. For the sake of transparency, permission to make available here the anonymous reviews motivating the rejection (<em>ReviewsPIDs4fileFormats.odt</em>) was asked, but was left without response. Some images (screendumps) and text result files from file identification tools tested are included. Further, a simple xquery command file (BaseX) for <em>fetch:content-type</em>()<em>, </em>used for getting MIME-types for files, is also provided.</p>
Table S27: Target and identified unknown organic micropollutants detected in surface water samples taken during heavy rain events
<p>In the following table, peak intensities of detected organic micropollutants in water samples are displayed.</p> <p>This data table is part of the appendix of Chapter 4 of the PhD thesis “Novel approaches to identify drivers of chemical stress in small rivers” by Liza-Marie Beckers prepared at RWTH Aachen University and at the Helmholtz Centre for Environmental Research-UFZ. In Chapter 4, precipitation-related pollutant patterns and indicator compounds during heavy rain events were identified in the Holtemme River by nontarget screening and cluster analysis. The table contains peak heights of organic micropollutants detected in water samples taken during heavy rain events in the Holtemme River (Saxony – Anhalt, Germany). The table is structured into the following columns: Compound name, use class of compound (e.g., pharmaceutical or pesticide), distinction between target or identified unknown compounds, mass-to-charge ratio (m/z), retention time (RT), assignment to a pattern identified by cluster analysis (i.e., “Base” or “Quick”), the probability of belonging to the assigned pattern as number between 0 and 1 as well as the peak height of the compound in each sample. The samples are indicated by "B" for "bottle" and a number from 1-16. The use class “NA” indicates that now major use class for this compound could be identified.</p> <p>The sampling was triggered by combined sewer overflow at a wastewater treatment plant upstream of the sampling point. Samples were taken by an automated sampler in 30-min composite samples for 8 hours resulting in 16 samples per rain event. In total, 6 heavy rain events from May to September 2016 were sampled during this study. The table is divided into 6 subtables (i.e., Table S27 A-F). Each subtable displays compounds and their peak heights detected in samples from one heavy rain event. The different rain events are abbreviated by the sampling date:</p> <p>Table S27A displays results from the rain event samples May 29<sup>th</sup> 2016 : E2905</p> <p>Table S27B displays results from the rain event samples June 01<sup>st</sup> 2016 : E0106</p> <p>Table S27C displays results from the rain event samples June 24<sup>th</sup> 2016 : E1306</p> <p>Table S27D displays results from the rain event samples June 13<sup>th</sup> 2016 : E2406</p> <p>Table S27E displays results from the rain event samples July 13<sup>th</sup> 2016 : E1307</p> <p>Table S27F displays results from the rain event samples September 17<sup>th</sup> 2016 : E1709</p> <p>Chemical analysis of the water samples was performed by liquid chromatography (UltiMate 3000 LC system (Thermo Scientific)) coupled to high resolution mass spectrometry (Q Exactive Plus, Thermo Scientific) with a heated electrospray ionization (HESI) source. Nontarget screening was performed as it allows for a comprehensive characterization of the chemical exposure during heavy rain events. However, only annotated target compounds and unknown compounds identified by structure elucidation are presented in the table. Details on data evaluation methods are described in Chapter 4 of the PhD thesis.</p> <p>Beckers, L.M. (2019): Novel approaches to identify drivers of chemical stress in small rivers. RWTH Aachen University, Aachen.</p>
Identifying Coronal Mass Ejection Active Region Sources: An automated approach - Catalogue results
<p>Catalogue of Coronal Mass Ejection (CME) active region sources. Includes a database version and a simplified .csv version. For full details, refer to the source code at <a href="https://github.com/JulioHC00/cmesrc">https://github.com/JulioHC00/cmesrc</a>. We include a README file for each describing each column.</p> <p>We also include the raw data used to generate the catalogue so that results may be reproduced following the steps detailed in <a href="https://github.com/JulioHC00/cmesrc">https://github.com/JulioHC00/cmesrc</a>. This is a collection of data from other works and we provide it only to allow the results to be reproduced</p> <p>Below, we detail the data sources for the raw_data folders</p> <p>==============================<br><strong>RAW DATA SOURCES</strong><br>==============================</p> <p><strong>DIMMINGS FOLDER</strong></p> <p>Data is from Solar Demon, .csv was provided by Emil Kraaikamp through private communication.</p> <blockquote> <p>Solar Demon – an approach to detecting flares, dimmings, and EUV waves on SDO/AIA images<br>Emil Kraaikamp, Cis Verbeeck<br>J. Space Weather Space Clim. 5 A18 (2015)<br>DOI: 10.1051/swsc/2015019</p> </blockquote> <p><strong>HARPNUM_TO_NOAA FOLDER</strong></p> <p>Obtained from http://jsoc.stanford.edu/doc/data/hmi/harpnum_to_noaa/all_harps_with_noaa_ars.txt</p> <p><strong>LASCO FOLDER</strong></p> <p>This CME catalog is generated and maintained at the CDAW Data Center by NASA and The Catholic University of America in cooperation with the Naval Research Laboratory. SOHO is a project of international cooperation between ESA and NASA.</p> <p>Downloaded from https://cdaw.gsfc.nasa.gov/CME_list/</p> <p><strong>MVTS FOLDER</strong></p> <p>Data from</p> <blockquote> <p>Angryk, R.A., Martens, P.C., Aydin, B. et al. Multivariate time series dataset for space weather data analytics. Sci Data 7, 227 (2020). https://doi.org/10.1038/s41597-020-0548-x</p> </blockquote> <p>Available at the Harvard Dataverse</p> <blockquote> <p>Angryk, Rafal; Martens, Petrus; Aydin, Berkay; Kempton, Dustin; Mahajan, Sushant; Basodi, Sunitha; Ahmadzadeh, Azim; Xumin Cai; Filali Boubrahimi, Soukaina; Hamdi, Shah Muhammad; Schuh, Micheal; Georgoulis, Manolis, 2020, "SWAN-SF", https://doi.org/10.7910/DVN/EBCFKM, Harvard Dataverse, V1</p> </blockquote> <p>The DT_SWAN folder contains the same data but with extra columns obtained directly from the Joint Science Operations Center (JSOC) through the python package drms.</p>
NMR and MS data of identified bromotyrosine alkaloids produced and released by Aplysina cavernicola
<p>This folder contains NMR and MS datasets of each identified bromotyrosine spiroisoxazoline pertaining to the publication<em> </em>entitled:</p> <p><strong>Diving into the molecular diversity of <em>Aplysina cavernicola’s </em>exo-metabolites: contribution of bromo-spiroisoxazoline alkaloids.</strong> <em>ACS Omega</em> 2022 <strong> <a href="https://pubs.acs.org/doi/10.1021/acsomega.2c05415"> </a></strong><a href="https://pubs.acs.org/doi/10.1021/acsomega.2c05415">https://doi.org/10.1021/acsomega.2c05415</a></p> <ul> <li>All NMR data were acquired in CD<sub>3</sub>OD at 600 MHz (Bruker Avance III, cryosonde TCI) using 2 mm NMR tubes</li> <li>All MS<sup>2</sup> data were acquired on a Bruker Impact II qTOF (ESI positive, collision energy 20-40eV)</li> </ul> <p>The compressed folder of the newly described Aplysine1 contains also raw data related to circular dichroism (CD) and infrared (IR) analyses, as well as quantum mechanical calculations of <sup>13</sup>C NMR shifts using GIAO NMR and DP4+ analyses.</p> <p>The Excel spreadsheet for DP4+ analyses were obtained from: Grimblat N et al. “Beyond DP4: An Improved Probability for the Stereochemical Assignment of Isomeric Compounds Using Quantum Chemical Calculations of NMR Shifts.” <em>The Journal of Organic Chemistry</em> 80, no. 24 (December 18, 2015): 12526–34. <a href="https://doi.org/10.1021/acs.joc.5b02396">https://doi.org/10.1021/acs.joc.5b02396</a>.</p> <p>All MS data are also made Freely available at the UCSD Center for Computational Mass Spectrometry database with the MassIVE identifier <a href="https://massive.ucsd.edu/ProteoSAFe/dataset.jsp?task=4f6d3c00539a412a9c6d7fac0f7f2a81">MSV000089502</a> .</p> <p>NOTE: 3,5 dibromotyrosine was not identified neither in<em> Aplysina cavernicola </em>crude extract nor as exo-metabolites but was used for MS dereplication purposes.</p>
Multi-omics identify LRRC15 as a COVID-19 severity predictor and persistent pro-thrombotic signals in convalescence
<p>RNA sequencing, SomaLogic proteomics and flow cytometry data were generated for two cohorts of end-stage kidney disease patients with COVID-19. The Wave 1 cohort consists of samples collected from patients during the first wave of COVID-19 in early 2020, while samples were collected for the Wave 2 cohort in the following year.</p> <p>This data deposition includes the RNA-seq counts, SomaScan proteomics, flow cytometry and clinical metadata associated with the study. For further information about the study and data, see the associated GitHub repository (https://github.com/jackgisby/covid-longitudinal-multi-omics) or our pre-print (https://doi.org/10.1101/2022.04.29.22274267). The repository also contains code to replicate our analysis of the data.</p> <p>The raw RNA-seq reads were processed using the nf-core RNA-seq v3.2 pipeline before htseq-count was used to generate a raw counts matrix, which is included in this deposition (<code>htseq_counts.csv</code>). Three files make up the proteomics data: <code>sample_technical_meta.csv</code>, <code>feature_meta.csv</code> and <code>soma_abundance.csv</code>. The first two files contain metadata columns for the samples and protein features, respectively. The final file includes the unprocessed protein abundance data. The files <code>general_panel.csv</code> and <code>t_cell_panel.csv</code> contain the flow cytometry data, split into the general and T-cell panels, respectively. Finally, clinical metadata is available for the two cohorts described in this study (<code>w1_metadata.csv</code>, <code>w2_metadata.csv</code>).</p> <p>The features in the clinical metadata include:</p> <table> <thead> <tr> <th>Column Name</th> <th>Data Type</th> <th>Description</th> </tr> </thead> <tbody> <tr> <td>sample_id</td> <td>Character</td> <td>Unique identifier for samples</td> </tr> <tr> <td>individual_id</td> <td>Character</td> <td>Unique identifier for individuals</td> </tr> <tr> <td>ethnicity</td> <td>Character</td> <td>The individual's ethnicity (asian, white, black or other)</td> </tr> <tr> <td>sex</td> <td>Character</td> <td>The individual's sex (M or F)</td> </tr> <tr> <td>calc_age</td> <td>Integer</td> <td>Age in years</td> </tr> <tr> <td>ihd</td> <td>Character</td> <td>Information on coronary heart disease</td> </tr> <tr> <td>previous_vte</td> <td>Character</td> <td>Whether individuals have had venous thromboembolism</td> </tr> <tr> <td>copd</td> <td>Character</td> <td>Whether individuals have chronic obstructive pulmonary disease</td> </tr> <tr> <td>diabetes</td> <td>Character</td> <td>Whether individuals have diabetes, and, if so, the type of diabetes</td> </tr> <tr> <td>smoking</td> <td>Character</td> <td>Smoking status</td> </tr> <tr> <td>cause_eskd</td> <td>Character</td> <td>Cause of ESKD</td> </tr> <tr> <td>WHO_severity</td> <td>Character</td> <td>The peak (WHO) severity for the patient over the disease course</td> </tr> <tr> <td>WHO_temp_severity</td> <td>Character</td> <td>The (WHO) severity at time of sampling</td> </tr> <tr> <td>fatal_disease</td> <td>Logical</td> <td>Whether the disease was fatal</td> </tr> <tr> <td>case_control</td> <td>Character</td> <td>Whether the individual was COVID-19 <code>POSITIVE</code> or <code>NEGATIVE</code> at time of sampling. Convalescent patients are denoted by the label <code>RECOVERY</code></td> </tr> <tr> <td>radiology_evidence_covid</td> <td>Character</td> <td>Evidence of COVID-19 from radiology</td> </tr> <tr> <td>time_from_first_symptoms</td> <td>Integer</td> <td>The number of days since the individual first experienced COVID symptoms at time of sampling</td> </tr> <tr> <td>time_from_first_positive_swab</td> <td>Integer</td> <td>The number of days since the individual's first positive swab was taken at time of sampling</td> </tr> </tbody> </table>
Pathways to enhance electrochemical CO2 reduction identified through direct pore-level modeling (data for figures)
<p>This is the data used to create the figures in the article "Pathways to enhance electrochemical CO2 reduction identified through direct pore-level modeling".</p> <p>Published in EES Catalysis</p> <p>DOI: 10.1039/d3ey00122a<br> Evan Johnson<br> Etienne Boutin<br> Shuo Liu<br> Sophia Haussener</p> <p><br> Additional notes are given in the "ReadMe.txt" file.</p>
TESS Confirmed and First Identified SuperWASP Variable Stars
<p>This catalog consists of the TESS-confirmed and First Identified SuperWASP Variable Stars of types $\delta$ Scuti, $\gamma$ Doradus, RR Lyrae, eclipsing binary systems with pulsating components, rotating variables, and others. This dataset is part of a short summary submitted to RNAAS entitled "Identifying SuperWASP Detected Candidate Variables with TESS" (Zhou, A.-Y., 2023 Research Notes of the AAS, vol.7). </p><p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.