Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
14,867
datasets available to search
ShareScore release 0.9.0
Dataset results
14,867 results for “determination”
Correlator data for determination of the I=1 pion-pion scattering amplitude and timelike pion form factor from Nf=2+1 lattice QCD
<p>Bootstrap samples of all correlation functions involved in the analysis of pion-pion scattering data and the timelike pion form factor described in "The I =1 pion-pion scattering amplitude and timelike pion form factor from N f = 2 + 1 lattice QCD". Additionally, an analysis file is provided for each ensemble which stored the analysis choices made in that work. These data are intended for use with the Jupyter notebook located in https://github.com/ebatz/jupan, which provides an interface. This notebook performs the entire analysis chain discussed in the above paper. </p>
Determination of rapeseed areas based on Sentinel-2 data
<p>Determination of rapeseed areas on the basis of Sentinel-2 data. The areas were determined using supervised classifications.</p> <p>Useful presentation:</p> <p>http://fabspace.pl/wp-content/uploads/2017/11/2_Klasyfikacja.pdf</p>
The molecular architecture of the yeast spindle pole body core determined by Bayesian integrative modeling
<p>This repository pertains to the molecular architecture of the yeast spindle pole body (SPB), the structural and functional equivalent of the metazoan centrosome. Data from in vivo FRET and yeast two-hybrid, along with SAXS, X-ray crystallography, and electron microscopy were integrated by a Bayesian structure modeling approach.</p> <p>For more information about how to reproduce this modeling, see the <a href="https://salilab.org/spb/">Sali lab website</a> or the README file.</p>
Critical Assessment of automated Structure Determination of Proteins by NMR
<p>The community-wide initiative "Critical Assessment of Automated Structure Determination of Proteins by NMR (<strong>CASD-NMR</strong>)" was launched in 2009 to to evaluate the ability of automated methods to produce 3D protein structures from NMR data that closely match structures manually determined by experts.</p> <p>This dataset includes all the experimental data made available to the participants of CASD-NMR in the two completed rounds of the initiative.</p> <p>Also refer to http://www-nmr.cabm.rutgers.edu/blindtest/blind.html for additional details, including first release date and link to each final PDB entry</p>
BaRTv1.0: an improved barley reference transcript dataset to determine accurate changes in the barley transcriptome using RNA-seq
<p>Background<br> Time consuming computational assembly and quantification of gene expression and splicing analysis from RNA-seq data vary considerably. Recent fast non-alignment tools such as Kallisto and Salmon overcome these problems, but these tools require a high quality, comprehensive reference transcripts dataset (RTD), which are rarely available in plants.</p> <p>Results<br> A high-quality, non-redundant barley gene RTD and database (Barley Reference Transcripts – BaRTv1.0) has been generated. BaRTv1.0, was constructed from a range of tissues, cultivars and abiotic treatments and transcripts assembled and aligned to the barley cv. Morex reference genome (Mascher et al., 2017). Full-length cDNAs from the barley variety Haruna nijo (Matsumoto et al., 2011) determined transcript coverage, and high-resolution RT-PCR validated alternatively spliced (AS) transcripts of 86 genes in five different organs and tissue. These methods were used as benchmarks to select an optimal barley RTD. BaRTv1.0-Quantification of Alternatively Spliced Isoforms (QUASI) was also made to overcome inaccurate quantification due to variation in 5’ and 3’ UTR ends of transcripts. BaRTv1.0-QUASI was used for accurate transcript quantification of RNA-seq data of five barley organs/tissues. This analysis identified 20,972 significant differentially expressed genes, 2,791 differentially alternatively spliced genes and 2,768 transcripts with differential transcript usage.</p> <p>Conclusion<br> A high confidence barley reference transcript dataset consisting of 60,444 genes with 177,240 transcripts has been generated. Compared to current barley transcripts, BaRTv1.0 transcripts are generally longer, have less fragmentation and improved gene models that are well supported by splice junction reads. Precise transcript quantification using BaRTv1.0 allows routine analysis of gene expression and AS.</p>
Replication Data for: Determination of Intrinsic Effective Fields and Microwave Polarizations by High-Resolution Spectroscopy of Single NV Center Spins
<p>Data repository for: <strong>Determination of Intrinsic Effective Fields and Microwave Polarizations by High-Resolution Spectroscopy of Single NV Center Spins</strong></p> <p><em>Data description.pdf</em> describes the uploaded data.<br> <em>Data.xlsx</em> is the data represented in the paper.<br> <em>Esrfit_Npeak.m</em>, <em>Esrfit_xN.m</em>, <em>GaussianFunc.m</em>, <em>Gaussian_xN_Func.m</em>, <em>Lorentz_Func.m</em>, <em>Lorentz_xN_Func.m</em>, <em>Rabifit_xN.m</em>, <em>Rabi_xN_Func.m</em>, <em>FourierTransformRabi.m</em> are Matlab code files to transform and fit the data.</p>
Nationally Determined Contributions (Latest submissions)
<p>Links to the latest (up to October 2021) national climate action plans under the Paris Agreement</p>
"Determinants of football players' valuation: a systematic review" datasets
<p>3 interdependent tables are available and are the materials used for a systematic review of the determinants of football players' valuation. </p> <p>- model_specifications is a table where each row represents one of the 111 model specifications included in the systematic review. Characteristics of the article from which specification was retrieved (title, year, authors, journal), attributes of the specification (sample size, population, econometric modeling, etc.), the significance levels of included variables, and the associated coefficients for significant variables. </p> <p>- model_specifications_dictionnary precise the names of the columns of the table model_specifications.</p> <p>- variables_definitions_and_classification is a table that details all the 471 variables used in the 29 articles analysed with definitions quoted from articles when possible and presents a classification of these variables into 6 categories and several subcategories. </p>
Dataset for 'Weld map tomography for determining local grain orientations from ultrasound'
<p>This dataset contains data files and Jupyter notebooks used to produce figures in the manuscript 'Weld map tomography for determining local grain orientations from ultrasound'.</p>
Data of Chinese treatment group for the research work "Disentangling material, social, and cognitive determinants of human behavior and belief".
<p>This repository contains data files of Chinese treatment group for the research work "Disentangling material, social, and cognitive determinants of human behavior and belief".</p>
Geographic range size and species morphology determines the organization of sponge host-guest interaction networks across tropical coral reefs (Raw data)
<p>Datasets for the analysis developed in the Article "<em><strong>Geographic range size and species morphology determines the organization of sponge host-guest interaction networks across tropical coral reefs</strong></em>". For more information, please refer to the original publication.</p> <p>Network_Structural_Index_&_SpogeTraits.csv <- Structural Index for the sponge-dwelling fauna network, sponge accumulated area and sponges’ morphology.</p> <p>NWTA_CoralReefs_Sponges_ interactions.csv <- Relationship between host sponges and guest fauna in the Northwester Atlantic coral reefs</p> <p>NWTA_CoralReefs_Sponge_reacords.csv <- Sponge species incidence records in the Northwester Atlantic coral reefs</p> <p>sponges_morphological_description.csv <- Sponge morphological standardization</p> <p>Network.html <- Interactive sponge-dwelling fauna network</p> <p>Enjoy!<br> </p>
Identification of factors determining the process of aggregation/agglomeration of metal oxide nanoparticles in a biological medium
<p>The model allows to identify factors determining the process of aggregation/agglomeration of metal oxide nanoparticles in a biological medium and to verify the importance of ion adsorption and protein adsorption in this process. </p> <p>Model confirms the significant effect of protein adsorption on the hydrodynamic diameter of metal oxide particles in the biological medium, and does not confirm the significant effect of ion adsorption in this process. It’s an example of modeling the properties of nanoparticles, where apart from the descriptors describing the structure of nanoparticles, there are also parameters characterizing the medium.</p>
Photodegradation of phenol (τOH) by TiO2-based nanophotocatalysts determined in line with the SAPNet methodology
<p>The independent variable (predictor) is the intensity of photoluminescence at 398 nm and with use of logistic regression model connects ability of photocatalytic degradation of phenol (endpoint) with this experimentally derived property. The equation goes as follow:<br> 4.32(±2.10) – 0.051(±0.03)(PL398)</p> <p>The developed model is included into the SAPNet workflow (Structure-Activity Prediction Network). In an additional step of SAPNet workflow developed model correlates the structure of a nanomaterial to the selected endpoint- photodegradation of phenol (τOH) by titanium dioxide synthesized in the presence of ionic liquids (IL). </p> <p>Each sample is described by surface area, the amount of nitrogen and carbon atoms, ionic liquid decomposition rate (ΔIL), molar ratio and the type of cations and anions, that influences photoluminescence </p>
Introducing MADYS: the Manifold Age Determination for Young Stars | Full model database
<p>Complete database of stellar and substellar evolutionary models employed in the Manifold Age Determination for Young Stars (MADYS). For a description of the tool, please refer to the main paper. For a detailed description of individual files, it is advised to use the ad-hoc functions provided within the published package.</p> <p>Bibliographic reference: arxiv:2206.02446</p> <p>GitHub repository: https://github.com/vsquicciarini/madys</p>
Diffraction data underpinning the structure of StayGold determined by X-ray crystallography (PDB code 8BXT)
<p>Raw diffraction data underpinning the crystal structure of StayGold fluorescent protein.</p> <p>This is the raw data underpinning PDB entry 8BXT.</p>
CRISPR/dCas9-mediated DNA demethylation screen identifies driver epigenetic determinants of colorectal cancer (Processed data)
<p><strong>Background:</strong> Promoter hypermethylation of tumour suppressor genes is frequently observed during the malignant transformation of colorectal cancer (CRC). However, whether this epigenetic mechanism is an actual driver of cancer or is a mere consequence of the carcinogenic process remains to be elucidated.</p> <p><strong>Results: </strong>In this work we performed an integrative multi -omic approach to identify gene candidates with strong correlations between DNA methylation and gene expression in human CRC samples and a set of 8 colon cancer cell lines. As a proof of concept, we combined recent CRISPR-Cas9 epigenome editing tools (dCas9-TET1, dCas9-TET-IM) with a custom arrayed gRNA library to modulate the DNA methylation status of 56 promoters previously linked with strong epigenetic repression in CRC, and we monitored the potential functional consequences of such DNA methylation loss by means of a high-content cell proliferation screen. Overall, the epigenetic modulation of most of these DNA methylated regions had a mild impact in the reactivation of gene expression and in the viability of cancer cells. Interestingly, we found that epigenetic reactivation of RSPO2 in the tumour context was associated with a significant impairment in cell proliferation in p53-/- cancer cell lines and further validation with human samples demonstrated that the epigenetic silencing of RSPO2 is a mid-late event in the adenoma to carcinoma sequence.</p> <p><strong>Conclusions: </strong>These results highlight the potential role of DNA methylation as a driver mechanism of CRC and open up the venue for the identification of novel therapeutic windows based on the epigenetic reactivation of certain tumour suppressor genes.</p>
State Water Project, Genetic Determination of Population of Origin 2011-2021
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Primary production estimates from 14C uptake (in situ), determined by the incorporation of inorganic carbon into particulate organic carbon (POC) due to photosynthesis at selected light levels from CCE LTER process cruises in the California Current System, 2006 - 2021 (ongoing).
Primary productivity samples of seawater are taken each day shortly before noon on the CTD rosette up-cast during the CCE Process crusies (since 2006, ongoing). Light penetration below the surface is estimated from the Secchi disk depth. Niskin bottles from depths with ambient light intensities corresponding to light levels simulated by on-deck incubators are identified and sampled. Primary production is estimated from 14C uptake using this simulated in situ technique (followed by filtering) by which the assimilation of dissolved inorganic carbon by phytoplankton yields a measure (in µg/L/day) of the rate of photosynthetic primary production (particulate organic carbon, POC) at selected light levels in the euphotic zone within the CCE study area.
Abundance and parasitoid infection dynamics of Guinardia delicatula on the Northeast U.S. Shelf from 2006 to 2022 determined by Imaging FlowCytobot.
These data include abundances of the diatom, Guinardia delicatula (= Rhizosolenia delicatula), on the Northeast U.S. Shelf from 2006 to 2022 as part of Long-Term Ecological Research (NES-LTER). Abundances are determined from Imaging FlowCytobot (IFCB) deployed in-situ at ~4m depth at the nearshore Martha’s Vineyard Coastal Observatory (MVCO) from 2006 to 2022 and in underway mode (sampling near-surface seawater) on 24 NOAA EcoMon survey cruises from 2013 to 2022. Abundances based on both human and machine learning image classification are provided. Total G. delicatula abundances are divided into two categories based on whether G. delicatula exhibited current or recent infection by the protistan parasitoid, Cryothecomonas aestivalis. Four data tables are provided with abundance values separated by sampling scheme (time series or survey cruise) and image classification approach (human or machine learning).
Abundance and biomass of Hemiaulus on the Northeast U.S. Shelf from 2013 to 2023 determined by Imaging FlowCytobot.
These data include abundance and carbon concentration of the diatom Hemiaulus on the Northeast U.S. Shelf during 81 research cruises from 2013 to 2023 as part of Long-Term Ecological Research (NES-LTER). Abundances are determined from Imaging FlowCytobot (IFCB) deployed in three different sampling schemes: underway mode (sampling near-surface seawater) on NOAA EcoMon, HAB Cyst, and AMAPPS broadscale survey cruises from 2013 to 2023; in underway mode (sampling near-surface seawater) on NES-LTER transect cruises from 2017 to 2023, and in discrete mode (CTD rosette discrete samples from depth) on NES-LTER transect cruises. Results are based on machine learning image classification, with one data table provided per sampling scheme (broadscale underway, transect underway, and transect discrete). Hemiaulus data are provided in abundance per milliliter and micrograms of carbon per liter.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.