Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Figure 4 in Biogeography, land snails and incomplete data sets: the case of three island groups in the Aegean Sea
Figure 4. UPGMA tree based on the similarity matrix obtained by applying Jaccard's index to the binary data matrix.
Figure 2 in Phylogenetic analysis and a time tree for a large drosophilid data set (Diptera: Drosophilidae)
Figure 2. Phylogenetic tree showing the reconstructed ancestral geographical distributions for extant and ancestral drosophilids estimated by the maximum-likelihood algorithm. Extant geographical distributions were retrieved from the Drosophila Stock Center or from the ZipcodeZoo database. See Table S2 for geographical distributions.
Figure 1 in Phylogenetic analysis and a time tree for a large drosophilid data set (Diptera: Drosophilidae)
Figure 1. Timescale for drosophilids based on a maximum-likelihood (ML) analysis using a concatenated alignment (9917 bp) of six protein-coding nuclear genes. Several monophyletic branches have been collapsed, indicating that all taxa within that taxonomic rank form a cluster. Support values above branches are bootstrap proportions performed on the ML tree; values less than 50 are not shown.
FIGURE 2. A phylogenetic network for a hypothetical data set. This network represents the relationships between four taxa, A-D in Exploring character conflict in molecular data*
FIGURE 2. A phylogenetic network for a hypothetical data set. This network represents the relationships between four taxa, A-D. The length of branch (a) is proportional to the strength of support for the relationship (A,C)(B,D). The length of branch (b) is proportional to the strength of support for the relationship (A,B)(C,D). In this example there is conflicting support for both of these arrangements, but more weight is given to (A,C)(B,D) than to (A,B)(C,D).
ChemBioSim Recalibration: Twelve Preprocessed ChEMBL Data Sets
<p><strong>Studying and mitigating the effects of data drifts on ML model performance at the example of chemical toxicity data</strong></p> <p><strong>Project description</strong></p> <p>Machine learning models are powerful tools for the prediction of molecular properties or the biological activity of chemical compounds. However, to make these models useful and applicable, the confidence in the predictions should also be specified. For that purpose, models may be integrated in a conformal prediction (CP) framework that adds a calibration step to estimate the confidence of the predictions. CP models offer the advantage of ensuring a predefined error rate, as long as the test and training sets are exchangeable.</p> <p>In cases where the test data presents a drift from the descriptor space of the training data, or where assay setups change, this assumption may not be fulfilled and the models are not guaranteed to be valid. </p> <p>In this study, the performance of internally valid CP models was evaluated upon application to either newer time-split data or to external data. More specifically, temporal data drifts were analysed based on time-splits of twelve toxicity-related datasets from the ChEMBL database. Moreover, models trained on publicly available data for liver toxicity and MNT in vivo were applied on proprietary data to evaluate the discrepancies. In general it was observed that the training and (holdout) test sets were not exchangeable in the studied set-ups, and the models were therefore not applicable (i.e. non-valid CP models).</p> <p>To recover the validity of the models on the holdout test set, a strategy for updating the calibration set with data more similar to the holdout set was investigated. Restored validity is the main requisite for applying the CP models with confidence. However, this comes at the cost of decreased model efficiency, as more predictions are identified as inconclusive.</p> <p> </p> <p><strong>Dataset</strong></p> <p>The uploaded file contains the ChEMBL data used in the work for the manuscript “Studying and mitigating the effects of data drifts on ML model performance at the example of chemical toxicity data”.</p> <p>Twelve preprocessed datasets containing molecule chembl ID, SMILES, binary activity (i.e. 1 if active, 0 if inactive), publication year, and CHEMBIO descriptors are available for the following ChEMBL endpoints, extracted from ChEMBL Version 26:</p> <ul> <li> <p>CHEMBL220: Acetylcholinesterase (human), 2673 compounds</p> </li> <li> <p>CHEMBL4078: Acetylcholinesterase (fish), 3811 compounds</p> </li> <li> <p>CHEMBL5763: Cholinesterase, 2755 compounds</p> </li> <li> <p>CHEMBL203: EGFR erbB1, 4059 compounds</p> </li> <li> <p>CHEMBL206: Estrogen receptor alpha, 1416 compounds</p> </li> <li> <p>CHEMBL279: VEGFR 2, 5174 compounds</p> </li> <li> <p>CHEMBL230: Cyclooxygenase-2, 2020 compounds</p> </li> <li> <p>CHEMBL340: Cytochrome P450 3A4, 3316 compounds</p> </li> <li> <p>CHEMBL240: HERG, 4976 compounds</p> </li> <li> <p>CHEMBL2039: Monoamine oxidase B, 2534 compounds</p> </li> <li> <p>CHEMBL222: Norepinephrine transporter, 1566 compounds</p> </li> <li> <p>CHEMBL228: Serotonin transporter, 2111 compounds</p> </li> </ul> <p> </p> <p><strong>Usage</strong></p> <p>This dataset can be used as input to run the notebooks available at </p> <p><a href="https://github.com/volkamerlab/CPRecalibration_manuscript_SI">https://github.com/volkamerlab/CPRecalibration_manuscript_SI</a></p> <ol> <li> <p>Clone the GitHub repository.</p> </li> <li> <p>Download the dataset provided here.</p> </li> <li> <p>Copy the dataset (don’t extract) into the data folder of the cloned GitHub repository.</p> </li> <li> <p>Follow the instructions on GitHub.</p> </li> </ol>
Figure 3. Maximum likelihood topologies. A, cytochrome oxidase 1 fragments. B, internal transcribed spacer fragment. C, combined data set. Bootstrap supports over 75 in Integrative taxonomy of Parasabella and Sabellomma (Sabellidae: Annelida) from Australia: description of new species, indication of cryptic diversity, and translocation of some species out of their natural distribution range
Figure 3. Maximum likelihood topologies. A, cytochrome oxidase 1 fragments. B, internal transcribed spacer fragment. C, combined data set. Bootstrap supports over 75% shown on nodes. Scale bar, average of nucleotide substitutions per site.
Data set for the paper "Temperatures and cooling rates recorded by the New Caledonia ophiolite: implications for cooling mechanisms in young forearc sequences"
<p>Mineral data set for the paper "Temperatures and cooling rates recorded by the New Caledonia ophiolite: implications for cooling mechanisms in young forearc sequences"</p>
Data set for the paper "Temperatures and cooling rates recorded by the New Caledonia ophiolite: implications for cooling mechanisms in young forearc sequences"
<p>Mineral data set for the paper "Temperatures and cooling rates recorded by the New Caledonia ophiolite: implications for cooling mechanisms in young forearc sequences"</p>
Data set for publication.
<p>Sequencing data.</p>
Figure 6 in A scolopocryptopid centipede (Chilopoda: Scolopendromorpha) from Mexican amber: synchrotron microtomography and phylogenetic placement using a combined morphological and molecular data set
Figure 6. Single shortest cladogram for six genes and morphology in combination (14 570 steps) under parameter set 3221. Numbers above branches are jackknife frequencies> 50%. Navajo rugs (as explained in Fig. 5) for the six parameter sets shown below branches.
Figure 5 in A scolopocryptopid centipede (Chilopoda: Scolopendromorpha) from Mexican amber: synchrotron microtomography and phylogenetic placement using a combined morphological and molecular data set
Figure 5. Single shortest cladogram for six genes in combination (14 465 steps) under parameter set 3221. Numbers above branches are jackknife frequencies> 50%. Navajo rugs below branches depict monophyly (black) or nonmonophyly (white) of clades under the six parameter sets shown at left; grey box indicates monophyly in some but not all shortest cladograms.
Figure 4 in A scolopocryptopid centipede (Chilopoda: Scolopendromorpha) from Mexican amber: synchrotron microtomography and phylogenetic placement using a combined morphological and molecular data set
Figure 4. Strict consensus of nine best-fit cladograms based on morphological data in Table 2 under implied weights (k = 2, 3, 4, 5, and 6). GC values> 50% shown above branches for concavity constant k = 3. Position of Scolopocryptops simojovelensis highlighted.
Figure 3 in A scolopocryptopid centipede (Chilopoda: Scolopendromorpha) from Mexican amber: synchrotron microtomography and phylogenetic placement using a combined morphological and molecular data set
Figure 3. Scolopocryptops simojovelensis sp. nov. Visualizations of synchrotron tomography data of holotype. A, B, dorsolateral and oblique anterodorsal views of head. C, ventral view of forcipules. D, lateral view of coxopleuron of left leg 23, anterior to left. Scale bars = 0.5 mm.
Figure 2 in A scolopocryptopid centipede (Chilopoda: Scolopendromorpha) from Mexican amber: synchrotron microtomography and phylogenetic placement using a combined morphological and molecular data set
Figure 2. Scolopocryptops simojovelensis sp. nov. Holotype AMNH Ch-SH7. A, nearly dorsal view of tergites 17–22; arrows on TT17 and 18 indicate complete paramedian sutures. B, dorsolateral view of tergites 17–20; inset shows anastomizing ridges parallel to posterior margin on tergite 19. C, dorsal view of segment 23, showing tergite (T23), coxopleural process (cp), dorsomedial spinose process (ds) and ventral spinose process (vs) of prefemur. D, dorsolateral view of leg pairs 21–23. Scale bars: A, B, D = 1 mm; C = 0.5 mm.
Figure 1 in A scolopocryptopid centipede (Chilopoda: Scolopendromorpha) from Mexican amber: synchrotron microtomography and phylogenetic placement using a combined morphological and molecular data set
Figure 1. Scolopocryptops simojovelensis sp. nov. Holotype AMNH Ch-SH7. A, dorsolateral view of complete specimen. B, distal part of right leg 20, showing tibial spur (ti) and tarsal spur (ta). C, dorsolateral view of cephalic plate and right antenna. D, distal part of tarsus and pretarsus of right leg 20, showing accessory spurs (ac). Scale bars: A = 5 mm; B = 0.5 mm; C = 1 mm; D = 0.1 mm.
Data Set: Software-Related Fatal Failures: An Empirical Exploration of RISKS Reports
<p>The data set used in the empirical exploration of RISKS reports. This data set acts as a basis for further research and investigation to categorize and analyze software-related fatal failures spanning more than 30 years.</p>
Synthetic data set to evaluate and benchmark the performance of multiple linear regression algorithms in Scikit-Learn and SANElib
<p>The datasets respresent different numbers of columns and rows to measure the scalability of linear regression algorihms in terms of columns and rows.</p>
Lampsilis siliquoidea and L. radiata seven microsatellite loci data set
<p>The data set corresponds to genotypes of individuals belonging to <i>Lampsilis siliquoidea</i> and <i>L. radiata </i>which are two closely related freshwater mussel species [Bivalvia: Unionidae]. Individual genotypes consist of seven microsatellite loci developed by Eackles and King 2002. Genotypes were used to asses population genetic structure above and below waterfalls in the lower Great Lakes (USA) and to investigate the degree of hybridization between these two species. </p>
Raw data set of pigeon body mass measurements
<p>1. Animal-borne logging devices are now commonly used to record and monitor the movements, physiology and behaviours of free-living animals. It is imperative that the impacts these devices have on the animals themselves is minimised.</p> <p>2. One important consideration is the interaction between the body mass of the animal, and the mass of the device.</p> <p>3. Using captive homing pigeons, we demonstrate that birds lose the equivalent amount of body mass compared to that of the logging device attached. With our experiments, we calculated that the compensatory mass loss because of the logging device equates to a total loss of 1,140 kJ of energy to the bird, over the 25-day period. This equates to 32% per day of their total daily energy budget.</p> <p>4. We suggest that practitioners of biologging give due consideration to the possibility of a device-induced decrease in body mass when making decisions regarding device size, and when considering the period of the time of the year at which devices are attached.</p> <p>5. It appears, based on the results of the present study, that device attachment is likely to be most disruptive during periods of regulated mass change, especially when periods of mass gain precede periods in which stored energy reserves are extensively utilised.</p> <p>6. These findings have significant consequences for anyone using biologging technology on both wild and captive volant animals. Further studies utilising captive birds are now needed to fully understand how context- and species-dependent physiological responses to externally attached devices are.</p>
Research Data Set for PhD thesis "Political Expression in Web Defacements"
<p>The software used for collection and processing is available at <a href="https://github.com/mkrzmr/Political-Expression-in-Web-Defacements-Crawler">https://github.com/mkrzmr/Political-Expression-in-Web-Defacements-Crawler</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.