Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,577
datasets available to search
ShareScore release 0.9.0
Dataset results
2,577 results for “Inference”
Data from: Inferring the evolution of reproductive isolation in a lineage of fossil threespine stickleback, Gasterosteus doryssus
<p>Darwin attributed the absence of species transitions in the fossil record to his hypothesis that speciation occurs within isolated habitat patches too geographically restricted to be captured by fossil sequences. Mayr's peripatric speciation model added that such speciation would be rapid, further explaining missing evidence of diversification. Indeed, Eldredge and Gould's original punctuated equilibrium model combined Darwin's conjecture, Mayr's model, and 124 years of unsuccessfully sampling the fossil record for transitions. Observing such divergence, however, could illustrate the tempo and mode of evolution during early speciation. Here, we investigate peripatric divergence in a Miocene stickleback fish, <em>Gasterosteus doryssus</em>. This lineage appeared and, over ~8,000 generations, evolved significant reduction of twelve of sixteen traits related to armor, swimming, and diet, relative to its ancestral population. This was greater morphological divergence than we observed between reproductively isolated, benthic-limnetic ecotypes of extant <em>Gasterosteus aculeatus</em>. Therefore, we infer that reproductive isolation was evolving. However, local extinction of low-armoured <em>G. doryssus</em> lineages shows how young isolate populations often disappear, supporting Darwin's explanation for missing evidence and revealing a mechanism behind morphological stasis. Exctinction may also account for limited sustained divergence within the stickleback species complex and help reconcile speciation rate variation observed across time scales.</p>
Inferred protein interactions between coronavirus and human proteins
<p>This repository contains the protein-protein interactions inferred by mimicINT (<a href="https://github.com/TAGC-NetworkBiology/mimicINT" target="_blank" rel="noopener">https://github.com/TAGC-NetworkBiology/mimicINT</a>) between the proteins of seven human coronaviruses (HCoV-229E, HCoV-HKU1, HCoV-NL63, HCoV-OC43, MERS-CoV, SARS-CoV and SARS-CoV-2) and substantial fraction of the human proteome. This dataset was generated in the context of the RiPCoN project (H2020-SC1-PHE-CORONAVIRUS-2020, <a href="https://cordis.europa.eu/project/id/101003633" target="_blank" rel="noopener">https://cordis.europa.eu/project/id/101003633</a>).</p>
TABLE 1 in Skeletal reconstruction of fossil vertebrates as a process of hypothesis testing and a source of anatomical and palaeobiological inferences
<p>TABLE 1. — Ratio of the length of phalanx I-1 to the length of metatarsal I in several ceratopsids. Length measurements were made from images with metatarsal I and phalanx I-1 in the same focal plane, using ImageJ (Schneider <i>et al.</i> 2012). The median ratio was used to determine that the expected length of phalanx I-1 for UALVP 42 was approximately 88.45% the length of phalanx I-1 in UALVP 16248. The digital model was scaled accordingly for the reconstruction.</p><table><thead><tr><th><b>Specimen</b></th><th><b>Taxon</b></th><th><b>Ratio</b></th></tr></thead><tbody><tr><th>AMNH 5351 cast (right foot)</th><td><i>Centrosaurus apertus</i> (Lambe, 1905)</td><td>1.021</td></tr><tr><th>CMN 8547</th><td>Indeterminate chasmosaurine</td><td>0.986</td></tr><tr><th>TMP 2002.076.0001</th><td>Indeterminate pachyrinosaurin</td><td>0.893</td></tr><tr><th>CMN 41357</th><td><i>Vagaceratops irvinensis</i> (Holmes Holmes, Forster, Ryan & Shepherd, 2001)</td><td>0.885</td></tr><tr><th>TMP 1989.097.0001</th><td><i>Styracosaurus albertensis</i> Lambe, 1913</td><td>0.883</td></tr><tr><th>AMNH 5351 cast (left foot)</th><td><i>Centrosaurus apertus</i></td><td>0.859</td></tr><tr><th>Median</th><td>–</td><td>0.889</td></tr></tbody></table>
Dataset: Inferring Inherent Optical Properties of Sea Ice Using 360-Degree Camera Radiance Measurements
<p>New types of compact 360-degree cameras have recently appeared on the consumer technology market. Some of these allow users to access raw imagery, offering sensor-level data that can be directly exploited for absolute light quantification. This paves the way for easy-to-use, inexpensive and accessible radiance cameras that can be operated in a wide range of natural environments. </p> <p>This dataset presents the angular radiance distributions measured with the Insta360 ONE 360-degree camera in sea ice. We report vertical profiles of the light field structure at two sites reprensentative of distinct sea ice types: High Arctic multi-year ice and Chaleur Bay (Quebec, Canada) landfast first-year ice. </p> <p>This repository contains the radiometric data stored in <strong>Hierarchical Data Format (HDF5, h5)</strong> under the following names: </p> <ul> <li><strong><a href="https://zenodo.org/api/records/14263256/draft/files/oden-08312018-imf-fluo.h5/content" target="_blank" rel="noopener noreferrer">oden-08312018-imf-fluo.h5</a></strong></li> <li><strong><a href="https://zenodo.org/api/records/14263256/draft/files/baiedeschaleurs-03232022-imf-fluo.h5/content" target="_blank" rel="noopener noreferrer">baiedeschaleurs-03232022-imf-fluo.h5</a></strong></li> </ul> <p>The High Arctic dataset (<strong>oden-08312018-imf-fluo.h5</strong>) contains only one station, while the Chaleur Bay (<strong>baiedeschaleurs-03232022-imf-fluo.h5</strong>) has four that can be accessed using these tags: "station_1", "station_2", "station_3", "station_4". The radiance measurements at each depth are reported as 2-dimensionals arrays with the azimuth directions (0-359°, 1° resolution) as columns and the zenith directions (0-180°, 1° resolution) as lines. The routines (coded in python) for the data processing can be found in the following <a href="https://github.com/RaphaelLarouche/radiance_camera_insta360/tree/master_v01" target="_blank" rel="noopener">Github repository</a> (master_v01) or the <a href="https://zenodo.org/records/4660994" target="_blank" rel="noopener">Zenodo stored version</a>. </p> <p>The methodologies to carefully calibrated the 360-degree camera for radiometry purpose are described in this <a href="https://doi.org/10.1364/AO.524122" target="_blank" rel="noopener">pulibcation</a> and the raw calibration data can be found in this Zenodo <a href="https://zenodo.org/records/10278731" target="_blank" rel="noopener">repository</a>. </p> <p>Additionnal information on the fieldwork and the data analysis are described in the <a href="https://doi.org/10.31223/X5V955" target="_blank" rel="noopener">preprint</a>.</p>
Inferring cosmology from gravitational waves using non-parametric detector-frame mass distribution: Data Release
<p>Dataset release accompanying Inferring cosmology from gravitational waves using non-parametric detector-frame mass distribution.</p>
Archival Datasets for SuperNova Artificial Inference by Lstm neural networks (SNAIL)
<p>The spectral-observation dataset (enclosed in the file archival_spec_observations.tar.gz) is comprised of 3091 observed spectra from 361 SNe Ia, largely contributed from CfA (Blondin et al. 2012), BSNIP (Silverman et al. 2012), CSP (Folatelli et al. 2013) and Supernova Polarimetry Program (Wang & Wheeler 2008; Cikota et al. 2019a; Yang et al. 2020).</p> <p>The spectral-template dataset (enclosed in the file archival_spec_templates.tar.gz) includes 361 spectral templates, each of them (covering -15 to +33d with wavelength from 3800 to 7200 A) was generated from the available spectroscopic observations of an individual SN via a LSTM neural network model.</p> <p>The auxiliary photometry dataset (enclosed in the file archival_phot_observations.tar.gz) provides the B & V light curves of these SNe (in total, 196 available SNe Ia), that were used to calibrate the synthetic B-V color of the observed spectra.</p> <p>In additional, the two master catalogs give the detailed information about the 361 SNe and their spectroscopic observations, respectively. </p> <p>These datasets are associated to the paper "Spectroscopic Studies of Type Ia Supernovae Using LSTM Neural Networks" (Hu et al. 2022, ApJ, accepted).</p>
Fig. 6. Bayesian inference tree for 5519 in First Record of Poecilobdella nanjingensis (Hirudinida: Arhynchobdellida: Hirudinidae) from Taiwan and its Molecular Phylogenetic Position within the Family
Fig. 6. Bayesian inference tree for 5519 bp alignment positions of nuclear 18S rRNA, 28S rRNA, mitochondrial cytochrome c oxidase subunit I, and 12S rRNA markers. Numbers on nodes indicate bootstrap values for maximum likelihood and Bayesian inference posterior probabilities.
Calibrating phylogenies assuming bifurcation or budding alters inferred macroevolutionary dynamics in a densely sampled phylogeny of bivalve families
<p>Analyses of evolutionary dynamics can be profoundly affected by age calibrations of phylogenetic nodes under different models of lineage branching. Most time-calibrated molecular phylogenies of extant taxa assume a purely bifurcating model, where nodes are calibrated using the daughter lineage with the older first occurrence in the fossil record. Lineages can also split via budding, in which a parent lineage persists following the origin of a daughter lineage, and nodes are calibrated using the age of the lineage with the younger first occurrence. Here, we use the extensive fossil record of bivalve molluscs for a large-scale empirical test of how the choice of branching model affects macroevolutionary analyses. We time-calibrated 91% of nodes in a phylogeny of 97 extant bivalve families using 86 calibration points ranging in age from 2.59 to 485 Ma. Allowing budding-based calibrations minimizes conflict between the tree topology and timing of evolutionary events in the fossil record, reducing the summed duration of inferred "ghost lineages," from 6.76 billion yrs (Gyr; bifurcating model) to 1.00 Gyr (budding model). Adding 31 extinct paraphyletic families – many major groups contain such extinct taxa – shifts deep splits further back in time and raises ghost-lineage totals to 7.86 Gyr (bifurcating) and 1.92 Gyr (budding), but more accurately reflects the time since separation of lineages. Lineage-through-time plots from phylogenetic data scaled under a bifurcating model of evolution push more inferred bivalve diversification into the Paleozoic, conflicting with other palaeontological evidence on the magnitude of the end-Paleozoic extinction and subsequent recovery, and strongly reduce the magnitude of the Cenozoic diversification of the group. Consideration of the hypothesized branching model within a given clade is essential when node-calibrating phylogenies, and for a major clade with a robust fossil record, an evolutionary model that allows budding and does not force bifurcations is the most appropriate one, and likely common for many other clades as well.</p>
Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities: discovery cohort meta data and parsed TCR repertoire data
<p>Meta data corresponding the the discovery cohort for the paper, "Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities" by Magdalena L Russell, Aisha Souquette, David M Levine, Stefan A Schattgen, E Kaitlynn Allen, Guillermina Kuan, Noah Simon, Angel Balmaseda, Aubree Gordon, Paul G Thomas, Frederick A Matsen IV, and Philip Bradley. These meta data include: </p> <p>(1) a file mapping the SNP data subject IDs to the TCR repertoire data subject IDs (gwas_id_mapping.tsv)<br> (2) a file including the PCAir PCs, self-reported ancestry, and genomic ancestry for each subject (all_pc_air.txt)<br> (3) a file including the PCAir variance explained by each PC (all_pc_air_variance.txt)<br> (3) a file including the SNP ID, chromosome, hg19 position, allele, rsid, and quality control metrics for each SNP in the SNP array (emerson_snp_rs_data.tsv)<br> (4) a file including IMGT genes and sequences used for parsing TCRB repertoire data (human_vj_allele_cdr3_nucseqs.tsv)<br> (5) a file including predicted TRBD2 allele genotypes for each subject (emerson_trbd2_alleles.tsv)<br> (6) Parsed TCRB repertoire data. These raw data were first published in Emerson et. al, <em>Nature Genetics </em>2017. (emerson_parsed_tcrb.tgz)</p> <p><strong>Corresponding discovery cohort raw TCR repertoire data is available here: </strong>https: //doi.org/10.21417/B7001Z (ImmuneACCESS database)<br> <strong>Corresponding discovery cohort SNP data is available here:</strong> https: //www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs001918.v1.p1 (The database of Genotypes and Phenotypes, accession number: phs001918)<br> <br> <strong>Software tools designed to work with these data are available here:</strong> https://github.com/phbradley/tcr-gwas</p>
Harnessing single cell RNA sequencing to identify dendritic cell types, characterize their biological states and infer their activation trajectory
<p><strong>Summary: </strong>Dendritic cells (DCs) orchestrate innate and adaptive immunity, by translating the sensing of distinct danger signals into the induction of different effector lymphocyte responses, to induce different defense mechanisms suited to face distinct types of threats. Hence, DCs are very plastic, which results from two key characteristics. First, DCs encompass distinct cell types specialized in different functions. Second, each DC type can undergo different activation states, fine-tuning its functions depending on its tissue microenvironment and the pathophysiological context, by adapting the output signals it delivers to the input signals it receives. Hence, to better understand DC biology and harness it in the clinic, we must determine which combinations of DC types and activation states mediate which functions, and how.<br> To decipher the nature, functions and regulation of DC types and their physiological activation states, one of the methods that can be harnessed most successfully is ex vivo single cell RNA sequencing (scRNAseq). However, for new users of this approach, determining which analytics strategy and computational tools to choose can be quite challenging, considering the rapid evolution and broad burgeoning of the field. In addition, awareness must be raised on the need for specific, robust and tractable strategies to annotate cells for cell type identity and activation states. It is also important to emphasize the necessity of examining whether similar cell activation trajectories are inferred by using different, complementary methods. In this chapter, we take these issues into account for providing a pipeline for scRNAseq analysis and illustrating it with a tutorial reanalyzing a public dataset of mononuclear phagocytes isolated from the lungs of naïve or tumor-bearing mice. We describe this pipeline step-by-step, including data quality controls, dimensionality reduction, cell clustering, cell cluster annotation, inference of the cell activation trajectories and investigation of the underpinning molecular regulation. It is accompanied with a more complete tutorial on Github. We anticipate that this method will be helpful for both wet lab and bioinformatics researchers interested in harnessing scRNAseq data for deciphering the biology of DCs or other cell types, and that it will contribute to establishing high standards in the field.</p> <p> </p> <p><strong>Data:</strong></p> <p>1. negative_cDC1_relative_signatures.csv : Negative signatures for performing Connectivity Map (cMAP) Analysis</p> <p>2. positive_cDC1_relative_signatures.csv : Positive signatures for performing Connectivity Map (cMAP) Analysis</p>
CrossDomainTypes4Py: A Python Dataset for Cross-Domain Evaluation of Type Inference Systems
<p>This dataset contains python repositories mined on GitHub on January 20, 2021. It allows a cross-domain evaluation of type inference systems. For this purpose, it consists of two sub-datasets, each containing only projects from the web or scientific calculation domain, respectively. Therefore we searched for projects with dependencies to either <a href="https://numpy.org/">NumPy</a> or <a href="https://flask.palletsprojects.com/en/2.0.x/">Flask</a>. Furthermore, only projects with dependencies to <a href="http://mypy-lang.org/">mypy</a> were considered, because this should ensure that at least parts of the projects have type annotations. These can be used later as ground truth. Further details about the dataset will be described in an upcoming paper, as soon as it is published it will be linked here.<br> The dataset consists of two files for the two sub-datasets. The web domain dataset contains 3129 repositories and the scientific calculation domain dataset contains 4783 repositories. The files have two columns with the URL to the GitHub repository and the used commit hash. Thus, it is possible to download the dataset using shell or python scripts, for example, the pipeline provided by <a href="https://github.com/saltudelft/many-types-4-py-dataset">ManyTypes4Py</a> can be used.<br> If repositories do not exist anymore or are private, you can contact us via the following email address: bernd.gruner@dlr.de. We have a backup of all repositories and will be happy to help you. </p>
Identifying strengths and weaknesses of methods for computational network inference from single cell RNA-seq data
<p>These data files contain single-cell RNA-sequencing expression data (expression_data.zip) and pseudotime files (pseudotime.zip) used to conduct comparisons of network inference methods on six published single-cell RNA-sequencing datasets. The resulting networks generated from the network inference methods are also uploaded here (normalized_inferred_networks.zip and imputed_inferred_networks.zip). Finally, the gold standard networks we used as ground truth to measure accuracy of the inferred networks are uploaded here (gold_standard_datasets.zip).</p>
Neural relational inference to learn long-range allosteric interactions in proteins from molecular dynamics simulations
<p>MD simulations used in the studies of the publication "<strong>Neural relational inference to learn long-range allosteric interactions in proteins from molecular dynamics simulations</strong>"</p>
Supplementary Data - Using landscape genomics to infer genomic regions involved in environmental adaptation of soybean genebank accessions
<p><strong>File: 50K_GenotypesEU_raw_UHOH_SoySNP50K.csv.tgz </strong></p> <p>Genotyping data of SoySNP50k SNP array of 170 European soybean varieties.</p> <p>The array includes 51.955 SNP markers.</p> <p>Genotypes of each variety are in columns and each row is a SNP marker. Naming of markers follows the annotation of the soybean genome.</p> <p><strong>File: EUvarieties_infos.csv </strong></p> <p>Description of European varieties</p> <p>Contains variety name, country of origin, EU region and maturity group assignment.</p> <p> </p> <p><strong>File: Supplementary_Data_Haupt_Schmid.xlsx</strong></p> <p>Additional data derived from data analysis. Description of data contained within file (Worksheet "Summary")</p> <p> </p>
Fluid Transport and Storage in the Cascadia Forearc inferred from Magnetotelluric Data
<p>The 3D electrical resistivity model derived from magnetotelluric data inversion. The model file is in netCDF format, which can be viewed using any standard netCDF plotting tools (e.g. <a href="https://www.giss.nasa.gov/tools/panoply/">Panoply</a>).</p>
Supplementary Data to *Robust adaptive distance functions for approximate Bayesian inference on outlier-corrupted data*
<p>Supplementary code and data to <strong>Robust adaptive distance functions for approximate Bayesian inference on outlier-corrupted data</strong> by <strong>Y. Schaelte et al., 2021</strong>.</p> <p>The archive contains a <strong>README.rst </strong>for information on what is where and how to execute the study and generate the figures. The underlying code without the data can be found at the repository https://github.com/yannikschaelte/study_abc_rad, of which this archive is a snapshot.</p> <p> </p>
Data from: Integrating tracking and resight data enables unbiased inferences about migratory connectivity and winter range survival from archival tags
<p>Archival geolocators have transformed the study of small, migratory organisms but analysis of data from these devices requires bias correction because tags are only recovered from individuals that survive and are re-captured at their tagging location. Data and code provided in this repository can be used to replicate the simulation and Painted Bunting case study results presented by Rushing et al. (2021) showing that integrating geolocator recovery data and mark–resight data enables unbiased estimates of both migratory connectivity between breeding and nonbreeding populations and region-specific survival probabilities for wintering locations.</p>
ManyTypes4TypeScript: A Comprehensive TypeScript Dataset for Sequence-Based Type Inference
<p>In this paper, we present ManyTypes4TypeScript, a very large corpus for training and evaluating machine-learning models for sequence-based type inference in TypeScript. The dataset includes over 9 million type annotations, across 13,953 projects and 539,571 files. The dataset is approximately 10x larger than analogous type inference datasets for Python, and is the largest available for TypeScript. We also provide API access to the dataset, which can be integrated into any tokenizer and used with any state-of-the-art sequence-based model. Finally, we provide analysis and performance results for state-of-the-art code-specific models, for baselining. ManyTypes4TypeScript is available on Huggingface and Zenodo.</p> <p>This dataset was collected on January 22, 2022 and deduplicated with Allamanis code deduplication tool.</p>
SEACells: Inference of transcriptional and epigenomic cellular states from single-cell genomics data
<p>Processed data for the manuscript "" available on bioRxiv at ""</p> <p>Data is available for the following samples</p> <ol> <li>CD34+ Multiome data : 2 replicates </li> <li>T-cell depleted bone marrow Multiome data: 2 replicates </li> </ol> <p> </p> <p>The following counts and fragments files are available for each replicate </p> <ol> <li><sample>_filtered_feature_bc_matrix.h5: Feature counts from CellRanger ARC</li> <li><sample>_atac_fragments.tsv.gz: ATAC fragments file from Cellranger ARC</li> <li><sample>_atac_fragments.tsv.gz.tbi: Index files for ATAC fragments file from Cellranger ARC</li> </ol> <p> </p> <p>The following scanpy anndata objects are also available</p> <ol> <li>cd34_multiome_rna.h5ad: Anndata object with normalized data, cell type annotations and clusters for the RNA modality of CD34+ hematopoietic stem and progenitor cells </li> <li>cd34_multiome_atac.h5ad: Anndata object with peak counts, cell type annotations and clusters for the ATAC modality of of CD34+ hematopoietic stem and progenitor cells.</li> <li>cd34_multiome_rna.h5ad: Anndata object with normalized data, cell type annotations and clusters for the RNA modality of T-cell depleted bone marrow dataset.</li> <li>bm_multiome_atac.h5ad: Anndata object with peak counts, cell type annotations and clusters for the ATAC modality of T-cell depleted bone marrow dataset.</li> </ol>
Data from 'Disparate inventories of hypoxia gene sets across corals align with inferred environmental resilience'
<p>Aquatic deoxygenation has been flagged as an overlooked but key factor driving mass bleaching-induced coral mortality as oxygen supplies lower to concentrations that can elicit an aerobic metabolic crisis i.e., hypoxia. Surprisingly little is known of the fundamental hypoxia responsive gene set inventory corals possess to respond to deoxygenation. It is unclear whether variation in gene copy number across species exist that potentially affect gene expression with subsequent differences in the effectiveness of a given stress response. Here, we used an ortholog-based meta-analysis to investigate how hypoxia gene inventories differed amongst coral species to assess putative copy number variation (CNV) across 24 coral protein sets from species with a sequenced genome that span corals from the robust and complex clade. We found approximately a third of the investigated genes exhibited copy number differences, and these differences were species-specific rather than the robust-complex split.</p> <p>Zipped folders of OrthoFinder results:</p> <p>'Results_Feb16' contains results including all 24 coral species from 7 genera (<em>Acropora, Pocillopora, Stylophora, Montastrea, Montipora, Obricella, Porites</em>).</p> <p>'gene_sets_acropora_acuminata_only' contains results including just one species per genera with <em>Acropora acuminata</em>.</p> <p>'gene_sets_acropora_cytherea_only' contains results including just one species per genera with <em>Acropora cytherea</em>.</p> <p>'gene_sets_acropora_digitifera_only' contains results including just one species per genera with <em>Acropora digitifera</em>.</p> <p> </p> <p>Results and Interpretations from these analyses are published open access here: <a href="https://doi.org/10.3389/fmars.2022.834332">https://doi.org/10.3389/fmars.2022.834332</a></p> <p>Full citation: Alderdice R, Hume BCC, Kühl M, Pernice M, Suggett DJ, Voolstra CR. Disparate inventories of hypoxia gene sets across corals align with inferred environmental resilience. Front Mar Sci. 2022;9. doi:10.3389/fmars.2022.834332</p> <p>Scripts are available here: <a href="https://zenodo.org/record/6396671#.YoYpoS8RoZg">https://github.com/didillysquat/alderdice_2021</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.