Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,078
datasets available to search
ShareScore release 0.7.1
Dataset results
4,078 results for βSARSβ
amel-github/sars-ani: Releasing new data fields in SARS-ANI dataset
<p>2022-06-20 - Release v1.1</p> <p>The original SARS-ANI dataset displayed common and scientific names of the animal host as found in the information source and/or inferred from the literature or expert knowledge.<br> Misspelled animal names and errors in taxonomy can lead to incorrect scientific conclusions and poor policy design. Moreover, harmonized host names can aid integrating other datasets (e.g. data on host biological traits, geographic distribution, or association with other pathogens).<br> Therefore, for each event, we programmatically performed taxonomic validation of the animal host name, using the R package taxize (Chamberlain et al. 2013). For more information on our validation process, see the R script <strong>sars_ani_validation.R.</strong></p> <p>Version 1.1. contains seven fields related to the identification of the animal host:</p> <ul> <li> <p>host_com_orig: Most specific designation of the animal host provided by the source(s), in English.</p> </li> <li> <p>host_sci_orig: Scientific name of the animal host as mentioned in the source(s) (scientific names are harmonized so that only the first letter of the genus is capitalized).</p> </li> <li> <p>host_com_res: Common name of the animal host, harmonized against the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/">NCBI</a>) taxonomic backbone.</p> </li> <li> <p>host_sci_res: Scientific name of the animal host (resolved to species or subspecies level), harmonized against the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/">NCBI</a>) taxonomic backbone.</p> </li> <li> <p>host_colloq: The colloquial name of the host, i.e. the name commonly used to identify the animal in non-specialist language (e.g. "tiger" for "Sumatran tiger").</p> </li> <li> <p>host_sci_spec_res: The scientific name of the host resolved to the species level.</p> </li> <li> <p>family: Animal family of the animal host.</p> </li> </ul>
Host removal database: Homo sapiens, Sars-Cov-2, PhiX174
<p>πΎ <strong>cleanup-db</strong></p> <p>Kraken2 database, built upon a viral sequence masked human reference from:</p> <ul> <li>Handley, Scott A. (2020). <strong>Virus+ Sequence Masked Human Reference Genome (hg19)</strong> (1.0) [Data set]. Zenodo. [<a href="https://zenodo.org/record/4116107">10.5281/zenodo.4116107</a>]</li> </ul> <p>but separating chromosomes as artificial taxa to allow for QC, and includes Sars-Cov-2 and PhiX 174</p> <p>πΎ <strong>gutcheck-db</strong></p> <p>A very small DB containg some common gut bacteria and Human and Murine mitochondrial genome:</p> <ul> <li><em>Akkermansia muciniphila</em></li> <li><em>Bacteroides fragilis</em></li> <li><em>Bifidobacterium longum</em></li> <li><em>Blautia obeum strain</em></li> <li><em>Escherichia coli</em></li> <li><em>Enterococcus faecium</em></li> <li><em>Prevotella copri</em></li> </ul> <p> </p> <p>See: <a href="https://github.com/telatin/cleanup">https://github.com/telatin/cleanup</a></p>
Coswara: A respiratory sounds and symptoms dataset for remote screening of SARS-CoV-2 infection
<p>Coswara is a dataset containing diverse set of respiratory sounds and rich meta-data from COVID-19 positive and Non-COVID subjects.</p>
X-ray diffraction data for SARS-CoV2 spike glycoprotein N-terminal heptad repeat domain + SARS-CoV2(QEYKKEKE)
<p>X-ray diffraction dataset for SARS-CoV2 spike glycoprotein N-terminal heptad repeat domain + SARS-CoV2(QEYKKEKE) collected at the AMX beamline (17-ID-1) at the National Synchrotron Lightsource II, Brookhaven National Laboratory, Upton, NY, USA.</p> <p>Final XDS.INP file generated by autoPROC.</p> <p>Serialized request document for vector collection from LSDC.</p> <p>KB mirrors</p> <p>Detector: EigerX9M (Si)</p> <p>Approx. photon flux at 13475eV: 4E12 ph/s</p> <p>Approx. beam size: 5 x 7 um</p>
SARS-CoV-2 mRNA vaccines induce persistent human germinal centre responses
<p>These are the<strong> processed</strong> BCR repertoire bulk sequencing data described in <a href="https://doi.org/10.1038/s41586-021-03738-2">Turner & O'Halloran et al., Nature, 2021</a> (Fig 3b-d; Extended Data Fig 3; Extended Data Table 6). The corresponding <strong>raw</strong> sequencing reads are available on SRA under <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA731610">BioProject PRJNA731610</a>.</p> <p><strong>Summary</strong>: Bulk-sorted total plasmablasts from PBMCs and germinal centre B cells at 4 weeks after primary immunization from 3 vaccinees who had no prior history of infection with SARS-CoV-2. </p> <p><strong>Code: </strong>Code along with Docker container for reproducing the NGS data-based figures and analyses in the published paper can be <a href="https://github.com/julianqz/wustl_published/tree/main/nature_2021">found on GitHub</a>.</p> <p><strong>Metadata file</strong>: WU368_turner_et_al_nature_2021_meta.tsv</p> <p>Abbreviations:</p> <ul> <li>LN = lymph node</li> <li>PB = plasmablast</li> <li>GC = germinal centre</li> <li>mAb = monoclonal antibody</li> </ul> <p><strong>BCR data file</strong>: WU368_turner_et_al_nature_2021_bcr.tsv.gz</p> <p>In addition to the processed bulk sequences, also included are the heavy chains of 37 mAbs that had been validated to be spike-binding and that were used together with the bulk sequences for clonal lineage inference. The mAbs are annotated as "mab" in the "seq_type" column.</p> <p><strong>BCR data column descriptions</strong></p> <p>The columns largely follow the <a href="https://changeo.readthedocs.io/en/stable/standard.html">AIRR-C Rearrangement format</a>. The main deviation is that CDR3s are used, as opposed to IMGT-defined "junctions". Non-standard columns are noted below.</p> <ul> <li>v_call_genotyped: V gene annotation reassigned after individualized genotyping by <a href="https://tigger.readthedocs.io/en/stable/">TIgGER</a></li> <li>germline_[vdj]_call: clonal consensus germline sequence reconstructed via <a href="https://changeo.readthedocs.io/en/stable/methods/germlines.html">`CreateGermlines.py --cloned` using Change-O</a></li> <li>isotype: IGH[ADEGM]</li> <li>cdr3: CDR3 nucleotide sequence</li> <li>cdr3_length: CDR3 nucleotide sequence length</li> <li>cdr3_aa: CDR3 amino acid sequence</li> <li>collapse_count: number of duplicate IMGT-aligned V(D)J sequences that were collapsed by <a href="https://alakazam.readthedocs.io/en/stable/topics/collapseDuplicates/">`alakazam::collapseDuplicates`</a></li> <li>donor: vaccinee</li> <li>sample: sample ID (arbitrary)</li> <li>timepoint: time point at which sample was collected</li> <li>tissue: tissue from which sample was collected</li> <li>sorting: FACS sorting</li> <li>seq_type: sequence type (mAb or bulk)</li> <li>nuc_RS_19_312: number of replacement and silent mutations between IMGT-numbered nucleotide positions 19-312 along IGHV sequences, calculated by <a href="https://shazam.readthedocs.io/en/stable/topics/calcObservedMutations/">`shazam::calcObservedMutations`</a></li> <li>nuc_denom_19_312: number of informative nucleotide positions for counting mutations, excluding non-A/T/G/C positions (such as "N", "-", ".")</li> <li>nuc_RS_freq_19_312: nucleotide-level mutation frequency (= nuc_RS_19_312 / nuc_denom_19_312)</li> </ul>
Molecular Dynamics Simulation of SARS-CoV-2 Spike Protein
<p>Trajectory data corresponding to the manuscript, tentatively titled "Distant Residues Modulate the Conformational Opening in SARS-CoV-2 Spike Protein"</p> <p>Authors: Dhiman Ray, Ly Le, Ioan Andricioaei</p> <p>Affiliation: University of California Irvine, USA</p> <p>Description: Multiple unbiased simulations of 40 ns were performed for the SARS-CoV-2 spike protein. Frames are saved at 50 ps interval. The initial structures were generated from umbrella sampling simulation starting from PDB ID: 6VSB and 6VXX. The index at the end of filename stands for the umbrella sampling window from which the trajectory was initiated. The indices are not continuous as not all the umbrella sampling windows were used to start trajectories. Additionally 3 trajectories, each of length 80 ns, are included for the closed, partially open and fully open state. The topology is provided as a PDB file ("spike_dry.pdb").</p> <p>The trajectories are for the spike head only structure obtained from the CHARMM-GUI Covid-19 archive. No solvent or ions are included in the trajectory or the topology.</p> <p>Update: Additional trajectories and PDB files for D614G mutant added. Each trajectory is 40 ns long. The PDB files are named 6VXX_mutant_dry.pdb and 6VSB_mutant_dry.pdb for the closed and partially open state.</p> <p>Pre-print available: https://doi.org/10.1101/2020.12.07.415596</p>
MD simulations of SARS-CoV-2 Spike Protein under static electric fields
<p>This dataset contains trajectories corresponding to all-atom MD simulations of segments of the SARS-CoV-2 Spike Protein, and in-silico mutations, under the influence of moderate external electric fields. The final structures of some of the simulations were used to perform in-silico docking with ACE2 receptor to evaluate the effect of comformational changes (docking was perform with PyDOCK).</p> <p>The file trajectories_6vsb_dt1ns.zip contains trajectories of simulations that were performed on a segment of the Protein Data Bank ID 6VSB comprising RBD, SD1 and SD2. The file trajectories_6m0j_dt1ns.zip correspond to the RBD in Protein Data Bank ID 6M0J. The file trajectories_in-silico_mutations_dt1ns.zip correspond to simulations performed on in-silico generated mutations following the mutations corresponding to WHO Variants of Concern UK, South Africa and Brazil. In all cases, simulations were performed at different electric field intensities ranging between 10<sup>4</sup> V/m and 10<sup>7</sup> V/m, with an extra short simulation under very high intensity (10<sup>9</sup> V/m). The file docked_structures_6m0j.zip contains the 100 best scored docked structures for each case as the output of PyDOCK.</p> <p>Trajectories are stored in GROMACS compressed trajectory file format (.xtc), downsampled to a 1ns timestep. Individual trajectories length are between 300 nanoseconds and 1 microsecond. In-silico docked structures are in PDB format. See linked preprint for more details.</p>
Banana Per Capita Consumption and SARS-CoV-2 Mortality Rates
<p>Plant Lectins are natural Antiviral agents and Banana is rich in them. Bananas are also a rich source of magnesium. Magnesium has a positive and effective role in increasing the body's immunity and stimulating the production of antibodies in the body.</p> <p>R2=0.99</p>
A vaccine-induced public antibody protects against SARS-CoV-2 and emerging variants
<p>These are the<strong> processed</strong> BCR repertoire bulk sequencing data described in <a href="https://doi.org/10.1016/j.immuni.2021.08.013">Schmitz, Turner & Liu et al., Immunity, 2021</a>. The <strong>raw</strong> sequence data are available on SRA under BioProjects <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA731610">PRJNA731610</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA741267">PRJNA741267</a>. </p> <p><strong>Summary</strong>: Bulk-sorted total plasmablasts and IgDlo enriched B cells from PBMCs and germinal centre B cells from lymph nodes from various timepoints after primary immunization from 22 BNT162b2 vaccinees who had no prior history of infection with SARS-CoV-2. </p> <p><strong>Metadata file</strong>: WU368_schmitz_et_al_immunity_2021_meta.tsv</p> <p>Abbreviations:</p> <ul> <li>LN = lymph node</li> <li>PB = plasmablast</li> <li>GC = germinal center</li> <li>mAb = monoclonal antibody</li> </ul> <p><strong>BCR data file</strong>: WU368_schmitz_et_al_immunity_2021_bcr.tsv.gz</p> <p>In addition to the processed bulk sequences, also included are the heavy chains of 37 mAbs (including 2C08) first reported in <a href="https://doi.org/10.1038/s41586-021-03738-2">Turner & O'Halloran et al., Nature, 2021</a> that had been validated to be spike-binding. The mAbs are annotated as "mab" in the "seq_type" column.</p> <p><strong>Sequence data column description</strong></p> <p>The columns largely follow the <a href="https://changeo.readthedocs.io/en/stable/standard.html">AIRR-C Rearrangement format</a>. The main deviation is that CDR3s are used, as opposed to IMGT-defined "junctions". Non-standard columns are noted below.</p> <ul> <li>v_call_genotyped: V gene annotation reassigned after individualized genotyping by <a href="https://tigger.readthedocs.io/en/stable/">TIgGER</a></li> <li>isotype: IGH[ADEGM]</li> <li>cdr3: CDR3 nucleotide sequence</li> <li>cdr3_length: CDR3 nucleotide sequence length</li> <li>cdr3_aa: CDR3 amino acid sequence</li> <li>donor: vaccinee ID</li> <li>sample: sample ID (arbitrary)</li> <li>timepoint: time point at which sample was collected</li> <li>tissue: tissue from which sample was collected</li> <li>sorting: FACS sorting</li> <li>seq_type: sequence type (mAb or bulk)</li> </ul>
Inflammatory Responses in the Placenta upon SARS-CoV-2 Infection Late in Pregnancy - IHC data
<p>SARS-CoV-2 infection during pregnancy does not affect the large majority of neonates but presents an increased risk for adverse pregnancy outcome. The effects of SARS-CoV-2 of its recently identified variants on placental function are not well understood. In this study, we investigated the impact of late gestational SARS-CoV-2 infection on the placenta.</p> <p>This dataset of is comprised of 897 images of classic immunohistochemistry for 3 markers in placenta from COVID-19 patients and controls.</p> <p><strong>A full description of the tissues, markers, and donors is available in the metadata.csv file.</strong></p>
Coarse-grained molecular dynamics simulations of SARS-CoV-2 envelope protein E in the pentameric form
<p>The trajectories of coarse-grained (CG) molecular dynamics (MD) simulations of<br> 1) unmodified (FeigLab_NMR; FeigLab_PentamerNoPTM_POPC_Martini3b: 5 μs; 5 μs); <br> 2) palmitoylated (FeigLab_PentamerCYSP43; PentamerCYSP44_POPC_Martini3b: 5 μs; 5 μs); <br> SARS-CoV-2 E protein pentamer in a POPC bilayer.</p> <p>The trajectory of CG MD of system containing 2 pentamers in the membrane buckled in a single direction (BuckledMembrane_FeigLab_2xPentamerNoPTM_POPC_Martini3b: 1 μs).</p> <p>FeigLab_Pentamer: https://github.com/feiglab/sars-cov-2-proteins/blob/master/Membrane/E_protein.pdb<br> FeigLab_NMR_Pentamer is assembled based on the transmembrane domain determined by NMR (PDB ID: 7K3G) and FeigLab model for the rest.</p>
Rutford Ice Stream, Antarctica M_sf tidal velocity components derived from COSMO-SkyMED SAR data
<p>This repository provides rasters for velocity components of a tidal (periodic) model for Rutford Ice Stream (RIS), Antarctica. The tidal model consists of a secular (constant) term and a single sinusoidal component corresponding to the M_sf tidal cycle (14.76529 days). The tidal model is fit to time-dependent velocity fields over RIS derived from speckle tracking of COSMO-SkyMed SAR data, collected over 9 months beginning in August 2013. The original methodology and source dataset are described in the publication:</p> <p>Minchew, B. M., Simons, M., Riel, B., & Milillo, P. (2017). Tidally induced variations in vertical and horizontal motion on Rutford Ice Stream, West Antarctica, inferred from remotely sensed observations. <em>Journal of Geophysical Research: Earth Surface</em>, <em>122</em>(1), 167-190. doi: <a href="https://doi.org/10.1002/2016JF003971">10.1002/2016JF003971</a></p> <p>The rasters are provided in GeoTIFF format in the Polar Stereographic South (EPSG: 3031) coordinate system. The velocity components are also referenced to Polar Stereographic South coordinates. The pixel spacing is 400 meters (in both the X- and Y-directions). The individual files are:</p> <ol> <li>vx_secular.tif: secular velocity in X-direction in meters/day.</li> <li>vy_secular.tif: secular velocity in Y-direction in meters/day.</li> <li>vx_amp.tif: M_sf velocity amplitude in X-direction in meters/day.</li> <li>vy_amp.tif: M_sf velocity amplitude in Y-direction in meters/day.</li> <li>vx_phase.tif: M_sf velocity phase delay in X-direction in days.</li> <li>vy_phase.tif: M_sf velocity phase delay in Y-direction in days.</li> </ol>
Experience of COVID-19 disease and fear of the SARS-CoV-2 virus among Polish students
<p>The deposited files contain a database related to the study of the fear of COVID-19 among Polish students and a code book. It is connected with the article titled <em>Experience of COVID-19 disease and fear of the SARS-CoV-2 virus among Polish students</em></p>
Head-to-head comparison of nasal and nasopharyngeal sampling using SARS-CoV-2 rapid antigen testing in Lesotho
<p>These are pseudo-anonymised data from the MISTRAL study: "Head-to-head comparison of nasal and nasopharyngeal sampling using SARS-CoV-2 rapid antigen testing in Lesotho". The data dictionary explains the data available in the dataset. Between December 2020 and September 2021, 2131 individuals with either COVID symptoms or contact with a COVID-positive case were included from two hospitals in Lesotho and had a valid PCR results.</p>
Tagged Twitter timelines for users reporting SARS-CoV-2 infections and related data
<p>Twitter data was collected through the Twitter API v2.0, specifically through the timeline endpoint. Details about the inference of SARS-CoV-2 self-reports and the tagging of the full timeline of each user can be found in the <a href="https://github.com/digitalepidemiologylab/content_changes_paper">GitHub repository</a>. The larger dataset ("preprocessed_data.csv") consists of a total of 8,534,171 tweets posted by 30,856 users from January 1, 2020 to October to September 30, 2021.</p> <p>The raw data from Twitter, including tweet and user IDs, has been removed or anonymized in order to comply with the EPFL guidelines for data sharing.</p> <p>In particular, the date of the tweets was removed, the text of the tweets, URLs and URL domains have been substituted with the "text", "<URL>" and "<URL_DOMAIN>" token respectively.</p> <p>User IDs have been substitued with new IDs in the [0, number of users] range (e.g. U0, U1, ...) .</p> <p>Tweet IDs have been substitued with new IDs in the [0, number of tweets] range (e.g. T0, T1, ...) .</p> <p>In addition to self-explanatory columns about topics, emotions, URL classification and symptoms we tagged, we also share the columns:</p> <ul> <li>pdate: date of the SARS-CoV-2 infection self-report for that user (adjusted with SUTime)</li> <li>effective_date: date of the tweet adjusted with SUTime, when the SUTime columns is available.</li> <li>rel_effective_day(week, month): days (weeks, months) computed with respect to the positivity date (i.e. "pdate" column). Negative numbers refer to tweets posted before the user reported a COVID-19 infection on Twitter.</li> </ul>
Dataset of "Exposure to airborne SARS-CoV-2 in four hospital wards and ICUs of Cyprus. A detailed study accounting for day-to-day operations and aerosol generating procedures."
<p>The authors highly appreciate being contacted if the data is to be used for any purpose.</p> <p>The following data set was used in the study entitled "<strong>Exposure to airborne </strong><strong>SARS-CoV-2 in four hospital wards and ICUs of Cyprus. A detailed study accounting for day-to-day operations and aerosol generating procedures.</strong>" and published in <em>Heliyon</em> Journal.</p> <p>This study characterized the transmission dynamics of airborne SARS-CoV-2 in normal and intensive care units. The data were collected over the period of 2020. In total, 165 and 62 air and environmental samples, respectively, were collected in four COVID-19 wards and ICUs in Cyprus and analyzed by RT-PCR. The comparison between RT-PCR and an alternative method for SARS-CoV-2 detection in air that provides comparable results but is less cumbersome and time demanding, is also given in the tab "Comparison with BELD".</p> <p>The data from sampling airborne SARS-CoV-2 using a MOUDI impactor are not included in this document but can be found in the supplement of the relevant publication.</p> <p>Please refer to the manuscript and its supplementary material for more information about how the data was collected. </p> <p> </p>
A evoluΓ§Γ£o do vΓrus SARS-COV-2 no Nordeste brasileiro.
<p>Os dados estão distribuídos da seguinte forma:</p> <p><em><strong>Região</strong></em> – Nome da região, no caso Nordeste;</p> <p><em><strong>Estado</strong></em> – Nome da UF;</p> <p><em><strong>Município</strong></em> – Nome do município;</p> <p><em><strong>Coduf</strong></em> – Código da UF;</p> <p><em><strong>Codmun</strong></em> – Código do município;</p> <p><em><strong>CodRegiaoSaude</strong></em> - Código da Região de Saúde;</p> <p><em><strong>NomeRegiaoSaude</strong></em> – Nome da Respectiva região de saúde;</p> <p><em><strong>Data</strong></em> – Data da coleta dos dados;</p> <p><em><strong>SemanaEpi</strong></em> – Número de semanas epidemológica;</p> <p><em><strong>PopulacaoTCU2019</strong></em> – População do Município;</p> <p><em><strong>CasosAcumulado</strong></em> – Casos acumulados no Município;</p> <p><em><strong>CasosNovos</strong></em> – Novos casos no Município;</p> <p><em><strong>ÓbitosAcumulado</strong></em> – Total de óbitos acumulados no município;</p> <p><em><strong>ÓbitosNovos</strong></em> - Total de novos óbitos no município;</p>
Glycosylated models for: The diversity of the glycan shield of sarbecoviruses closely related to SARS-CoV-2
<p>Glycosylated models (as PDB files) of the sarbecovirus spike proteins used in the study: The diversity of the glycan shield of sarbecoviruses closely related to SARS-CoV-2.</p>
Supporting data for "CoVEffect: Interactive System for Mining the Effects of SARS-CoV-2 Mutations and Variants Based on Deep Learning"
<p>This repository contains the datasets created and extracted for the paper:</p> <p>Giuseppe Serna García, Ruba Al Khalaf, Francesco Invernici, Stefano Ceri, and Anna Bernasconi. 2022.<br> "<strong>CoVEffect</strong>: Interactive System for Mining the <strong>Effects of SARS-CoV-2 Mutations and Variants</strong> Based on Deep Learning". (Available online at http://gmql.eu/coveffect)</p> <p>--------------------------------------------------------------------------------<br> LIST OF FILES WITH DESCRIPTION:<br> --------------------------------------------------------------------------------</p> <p>AdditionalFile1-effects-taxonomy:<br> Descriptions of legal values for the 'Effect' field, based on a categorized taxonomy.</p> <p>AdditionalFile2-levels-taxonomy:<br> Descriptions of legal values for the 'Level' field.</p> <p>AdditionalFile3-training_dataset_target:<br> List of target tuples (manually annotated) of 221 abstracts considered for training the model. For each abstract, target tuples follow the schema ID, DOI, title, entity, effect, level, type (mutation or variant), tuples_count (>1 when an effect/level is shared by multiple entities, #abstracts containing the same effect described in the tuple).</p> <p>AdditionalFile4-validation_dataset_target:<br> List of target tuples (manually annotated) of 50 abstracts considered for validating the prepared prediction model.<br> For each abstract, target tuples follow the schema defined for AdditionalFile3.</p> <p>AdditionalFile5-validation_dataset_highlighted:<br> Textual abstracts of the 50 manuscripts considered for validation; the text used to support the manual target annotations has been highlighted in yellow.</p> <p>AdditionalFile6-validation_dataset_prediction:<br> List of predicted annotations of 50 abstracts considered for validating the prepared prediction model. The file is split in 4 TSV, respectively for entity (a), effect (b), level (c), and whole tuple predictions (d).</p> <p>AdditionalFile7-keywords_query_list:<br> Keyword-based search run on the CORD-19 dataset to extract a relevant subset of abstracts regarding the scope of interest of CoVEffect. The Boolean logic used to combine keywords is explained in the section 'Annotations of the biology-related CORD-19 cluster'.</p> <p>AdditionalFile8-CORD-19_batch_dataset_metadata:<br> Metadata of the 7,230 papers extracted by the keyword-based query in AdditionalFile7.<br> These abstracts have been annotated by the prediction framework.</p> <p>AdditionalFile9-CORD-19_batch_dataset_prediction:<br> List of predicted annotations of 7,230 abstracts extracted from the biology-related cluster of CORD-19.</p> <p>AdditionalFile10-test_dataset_target:<br> List of target tuples (manually annotated) of 100 abstracts randomly selected from the 7,230 extracted as in AdditionalFile8.<br> For each abstract, target tuples follow the schema defined for AdditionalFile3.</p> <p>AdditionalFile11-test_dataset_prediction:<br> List of predicted annotations of 100 abstracts considered for testing the prediction model on a subset of the CORD-19 biology-related cluster. As AdditionalFile6, it is split in 4 TSV, respectively for entity (a), effect (b), level (c), and whole tuple predictions (d).</p>
SARS-CoV-2 Infection and Clinical Signs in Cats and Dogs from Confirmed Positive Households in Germany
<p>Supplemental material and raw data referring to specified publication</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.