Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,655
datasets available to search
ShareScore release 0.9.0
Dataset results
3,655 results for “Structural data”
Bioactive compounds with no structural analogs (high-confidence activity data)
<p>A set of 52,815 unique bioactive compounds (human targets, high-confidence activity data) with no structural analogs with high-confidence activity data was extracted from ChEMBL. For each compound the ChEMBL compound ID (CHEMBLID_Compound) and high-confidence target annotation(s) (CHEMBLID_Targets) are provided. The data set was generated as a part of an analysis to be published in 'Medicinal Chemistry Communications'. </p>
Experimental data, Periodic responses of a structure with 3:1 internal resonance
<p>Dataset for</p> <p>Alexander D Shaw; Thomas L Hill; Simon A Neild; Michael I Friswell</p> <p>Periodic responses of a structure with 3:1 internal resonance</p> <p>MSSP, in press (as of 18/3/2016)</p> <p>10.1016/j.ymssp.2016.03.008</p> <p> </p>
LiDAR-derived forest structure data and predictions of the locations of old-growth forests for Central Finland.
<p><strong>INTRO</strong><br> This archive contains data and analysis code for the Biodiversity Map -project conducted by Open Knowledge Finland (http://fi.okfn.org/projects/biodiversity-map/)</p> <p><strong>LICENCE</strong><br> The files listed below are all released to the public domain under a CC0 public domain dedication (https://creativecommons.org/publicdomain/zero/1.0/)</p> <p><strong>FILE DESCRIPTIONS</strong></p> <p><em><strong>FILE 1:</strong></em> background.zip<br> Inside the archive is a comma-separated file "background.csv" containing LiDAR-derived forest structure variables for 2/3 of Central Finland. These were derived from 3 raster data sets describing forest canopy maximum height (mh), forest canopy cover (cc) and lidar return intensity (in). The rasters had resolutions of 6 metres, 6 metres and 2 metres, respectfully. An 18 m resolution grid was then used to aggregate the rasters into average, minimum and maximum values + standard deviations of the original variables. The original LiDAR data was made available by the National Land Survey of Finland.</p> <p><br> <em><strong>FILE 2:</strong></em> conservation.lambdas<br> This file contains fitted parameters for the maxent model. For more information, check maxent documentation at https://www.cs.princeton.edu/~schapire/maxent/</p> <p><strong><em>FILE 3:</em></strong> conserved_swd.csv<br> Forest structure variables at 18 meter resolution for old-growth conservation areas in Central Finland. A subset of background.csv. This file still has a header, the variables are the same as in background.csv</p> <p><em><strong>FILE 4:</strong></em> grass_create_forest_rasters_from_las.sh<br> A shell script used to convert LiDAR files to raster maps of forest structure with GRASS 7.</p> <p><em><strong>FILE 5:</strong></em> lidar_coverage.png<br> A map showing the extent of LiDAR data available for Central Finland when we did the analyses.</p> <p><em><strong>FILE 6:</strong></em> maxent_model_run_product.sh<br> A shell script used to fit the maximum entropy model to predict the locations of conservation-area-like forests in Central Finland.</p> <p><em><strong>FILE 7:</strong></em> projection_product.csv<br> The results of the maxent model in a comma separated file. The first row has the variable names: x,y,product_fit. x and y are coordinates in the CRS ETRS-TM35FIN (EPSG:3067). product_fit is "the probablility that this 18*18 meter grid cell is old-growth conservation area".</p> <p><em><strong>FILE 8:</strong></em> README<br> A file with a description of the dataset in human-readable form.</p> <p><strong>VALIDATION FILES</strong><br> The data in these files was collected to validate the results of the aforementioned maxent model. The data were collected in a hierarchical sampling scheme: six randomly determinded unintersecting 9 km * 9 km landscape windows were chosen for sampling. From each window, three samples were taken. One sample from conservation areas, one sample from the "best" 10 % of forests as determined by the maxent model excluding conservation areas and one random sample. Not all windows contained conservation areas, and not all areas were accessible (islands, for example). In addition a few areas were skipped due to time constraints.</p> <p>The sampled points are identified by their lanscape window (suuralue), their sample (otos) and their sample number (mittauspiste).</p> <p><em><strong>FILE 9:</strong></em> validation_felled.csv<br> A comma separated list of those points that were not measured because they were felled.</p> <p><em><strong>FILE 10:</strong></em> validation_gps_results_2016-09-07.csv<br> A list of gps coordinates for all the sample points. product_fit is the value of the geographically closest prediction from the maxent model described above.</p> <p><em><strong>FILE 11:</strong></em> validation_lying_deadwood_transects_2016-08-30.csv<br> A comma separated file with data from deadwood transects. From each validation point, three 30 m long transects were made with 120 degree angles between them, and all lying deadwood more than 2 cm in diameter were measured. For some validation points, there were geographical obstructions which prevented the full 90 m of transect being surveyed, this is also recorded in the data. Each row holds measurements from one lying trunk.<br> </p> <p><em><strong>FILE 12:</strong></em> validation_relascope_2016-08-30.csv<br> Relascope measurements from the validation points. Each row is measurements for one species from one validation point. Dead and alive trees are counted separately.<br> </p> <p><strong>MORE INFORMATION</strong></p> <p>For more in-depth descritions of the files, read the file named README.<br> For some auxilliary files and information, check our old hackathon repository on github: https://github.com/Koalha/bdm_hackathon</p>
Research data supporting "Sequence-Dependent Self-Assembly and Structural Diversity of Islet Amyloid Polypeptide-Derived β-Sheet Fibrils"
<p>Research data supporting the publication:</p> <p>Wang, S.-T. et al., 2017, Sequence-Dependent Self-Assembly and Structural Diversity of Islet Amyloid Polypeptide-Derived β-Sheet Fibrils, ACS Nano, http://dx.doi.org/10.1021/acsnano.7b02325</p>
Human ancestral structure data from cobraa
<p>This is an updated version for the data from our paper on human ancestral population structure. The previous upload has truncated marginal decoding files. Here, I upload the full decoding files from cobraa-path (each state is a tuple of time and path), from which the marginal path probabilities can be easily obtained (see below). I also upload the final inference files from cobraa, for panmictic (PSMC) inference and the best fitting structured inference. These files exist for all of the 26 populations in the 1000 Genomes Project (one sample per population).</p> <p>To get the marginal path probabilities, the script marginalise_fulldecoding.py can be used. Example usage (the file paths will have to be changed):<br>Usage<br>python human_ancestral_structure_v2/marginalise_fulldecoding.py -chrom 20 -popsam GBR_HG00118 -outprefix /home/trevor/testingdelete250531 -decode_file human_ancestral_structure/decoding/GBR_HG00118/chr20.txt.gz</p> <p>Write all in a bash loop with<br>for chrom in {1..22}; do for popsam in GBR_HG00118 TSI_NA20752 IBS_HG01783 FIN_HG00266 CEU_NA12718 CHS_HG00443 KHV_HG02113 CHB_NA18530 CDX_HG02373 JPT_NA18939 BEB_HG03006 PJL_HG03234 GIH_NA20845 STU_HG03753 ITU_HG03977 PUR_HG01171 CLM_HG01250 PEL_HG02285 MXL_NA19648 ESN_HG03515 YRI_NA18488 MSL_HG03212 GWD_HG02568 ACB_HG01882 ASW_NA19625 LWK_NA19017; do echo popsam=${popsam}, chrom=${chrom}; python human_ancestral_structure_v2/marginalise_fulldecoding.py -chrom ${chrom} -popsam ${popsam} -outprefix /home/trevor/testingdelete250531 -decode_file human_ancestral_structure/decoding/${popsam}/chr${chrom}.txt.gz ; echo; done; done</p> <p>Please post questions on the GitHub https://github.com/trevorcousins/cobraa</p>
Data from: warming and top-down control of stage-structured prey: linking theory to patterns in natural systems
<p>Warming has broad and often nonlinear impacts on organismal physiology and traits, allowing it to impact species interactions like predation through a variety of pathways that may be difficult to predict. Predictions are commonly based on short-term experiments and models, and these studies often yield conflicting results depending on the environmental context, spatiotemporal scale, and the predator and prey species considered. Thus, the accuracy of predicted changes in interaction strength, and their importance to the broader ecosystems they take place in, remain unclear. Here, we attempted to link one such set of predictions generated using theory, modeling, and controlled experiments to patterns in the natural abundance of prey across a broad thermal gradient. To do so, we first predicted how warming will impact a stage-structured predator-prey interaction in riverine rock pools between Pantala spp. dragonfly nymph predators and Aedes atropalpus mosquito larval prey. We then described temperature variation across a set of hundreds of riverine rock pools (n = 775) and leveraged this natural gradient to look for evidence for or against our model's predictions. Our model's predictions suggested that warming should weaken predator control of mosquito larval prey by accelerating their development and shrinking the window of time that aquatic dragonfly nymphs could consume them in. This was consistent with data collected in rock pool ecosystems, where the negative effects of dragonfly nymph predators on mosquito larval abundance were weaker in warmer pools. Our findings provide additional evidence to substantiate our model-derived predictions, while emphasizing the importance of assessing similar predictions using natural gradients of temperature whenever possible.</p>
Data Set for the Journal Article "Automated Preparation of Nanoscopic Structures: Graph-Based Sequence Analysis, Mismatch Detection, and pH-Consistent Protonation with Uncertainty Estimates"
<p>This repository containes the data generated by ASAP and discussed in the journal article [Csizi, K.-S. and Reiher, M., 2023, arXiv:2307.16344], including Cartesian coordinates of training and test set molecules, and MD trajectories. </p>
Data and code corresponding to the article "Interaction network structure explains species temporal persistence in empirical plant-pollinator communities"
<p>This upload contains the Datasets and code to generate the results of the article "Interaction network structure explains species temporal persistence in empirical plant-pollinator communities".</p><p>The database comprises two files containing the abundances of plants and pollinators, and one containing the interaction networks among plants and pollinators. </p><p>The code folder contains the code to generate the results, and to generate the figures of the manuscript. </p>
Research data supporting: "Machine learning of microscopic structure-dynamics relationships in complex molecular systems"
<p>This repository contains the set of data and the code to reproduce the results shown in "Machine learning of microscopic structure-dynamics relationships in complex molecular systems" published on Machine Learning: Science and Technology (DOI: 10.1088/2632-2153/ad0fa5).</p>
Data for: Genetic structuring and species boundaries in the Atlantic stony coral Favia (Scleractinia, Faviidae)
<p class="MsoNormal">Scleractinian corals are the main modern builders of coral reefs, dynamic ecosystems that are hot spots of marine biodiversity. Southern Atlantic reef corals are understudied compared to their Caribbean and Indo-Pacific counterparts and many hypotheses about their population dynamics demand further testing. We employed thousands of single nucleotide polymorphisms (SNPs) recovered via ezRAD to characterize genetic population structuring and species boundaries in the amphi-Atlantic hard coral genus <em>Favia</em>. Coalescent-based species delimitation (BFD* - Bayes factor delimitation) recovered <em>F. fragum </em>and <em>F. gravida </em>as separate species. Although our results agree with depth-related genetic structuring in <em>F.</em><em> frag</em><em>um</em><em>,</em><em> </em>they did not support incipient speciation of the "tall" and "short" morphotypes. The preferred scenario revealed a split between two main lineages of <em>F. gravida</em>, one from Ascension Island and the other from Brazil. The Brazilian lineage is further divided into a species that occurs throughout the Northeastern coast and another that ranges from the Abrolhos Archipelago to the state of Espírito Santo. BFD* scenarios were supported by analysis of datasets with varying levels of missing data. Our results challenge current notions about Atlantic reef corals because they uncovered surprising genetic diversity in <em>Favia</em><em> </em>and<em> </em>rejected the long-standing hypothesis that Abrolhos Archipelago may have served as a Pleistocenic refuge during the last glaciations. </p>
Experimental data for "Exact inversion of partially coherent dynamical electron scattering for picometric structure retrieval"
Open the record for dataset details and reuse information.
Matlab code to calibrate a structured-PDE model to data from in vitro experiments
<p>Matlab code for the calibration of a PDE model of evolutionary dynamics of a well-mixed population of aggressive breast cancer cells from in vitro data on MCF7-sh-WISP2 cell line and bootstrapping for uncertainty quantification. For more details, see the associated publication: "Evolutionary dynamics of glucose-deprived cancer cells: insights from experimentally-informed mathematical modelling", by L. Almeida, J. Denis, N. Ferrand, T. Lorenzi, A. Prunet, M. Sabbah, C. Villa (corresponding author, author of code), 2023. In press in the journal of the Royal Society Interface.</p>
Data from: The role of fish predators and their foraging traits in shaping zooplankton community structure
<p><span>Differentiation of foraging traits among predator populations may help explain observed variation in the structure of prey communities. However, few studies have investigated the phenotypic effects of predators on their prey in natural communities. Here, we use a comparative analysis of 78 Greenlandic lakes to examine how foraging trait variation among threespine stickleback populations can help explain variation in zooplankton community composition among lakes. We find that landscape-scale variation in zooplankton composition was jointly explained by lake properties, such as size and water chemistry, and the presence and absence of both stickleback and arctic char. </span><span>Additional variation in zooplankton community structure can be explained by stickleback jaw protrusion, a trait with known utility for foraging on zooplankton, but only in lakes where stickleback co-occur with arctic char. Overall, our results illustrate how trait variation of consumers, alongside other ecosystem properties, can influence the composition of prey communities in nature.</span></p>
Data for "Influence of variation in grain boundary parameters on the evolution of atomic structure and properties of [111] tilt grain boundaries in aluminum"
<p>This repository contains the raw data of experimental STEM images and of the simulations for the paper "Influence of variation in grain boundary parameters on the evolution of atomic structure and properties of [111] tilt boundaries in aluminum".</p>
Data from: Social complexity affects cognitive abilities but not brain structure in a Poecilid fish
<p>Some cognitive abilities are suggested to be the result of a complex social life, allowing individuals to achieve higher fitness through advanced strategies. However, most evidence is correlative. Here, we provide an experimental investigation of how group size and composition affect brain and cognitive development in the guppy (<em>Poecilia reticulata</em>). For six months, we reared sexually mature females in one of three social treatments: a small conspecific group of three guppies, a large heterospecific group of three guppies and three splash tetras (<em>Copella arnoldi</em>) – a species that co-occurs with the guppy in the wild, and a large conspecific group of six guppies. We then tested the guppies' performance in self-control (inhibitory control), operant conditioning (associative learning), and cognitive flexibility (reversal learning) tasks. Using X-ray imaging, we measured their brain size and major brain regions. Larger groups of six individuals, both conspecific and heterospecific groups, showed better cognitive flexibility than smaller groups, but no difference in self-control and operant conditioning tests. Interestingly, while social manipulation had no significant effect on brain morphology, relatively larger telencephalons were associated with better cognitive flexibility. This suggests alternative mechanisms beyond brain region size enabled greater cognitive flexibility in individuals from larger groups. Although there is no clear evidence for the impact on brain morphology, our research shows that living in larger social groups can enhance cognitive flexibility. This indicates that the social environment plays a role in the cognitive development of guppies.</p>
Data for the publication: High performance of porous, hierarchically structured P2- Na0.6Al0.11 – xNi0.22 – yFex+yMn0.66O2 cathode materials
<p>Data sets: SEM-images, EIS, ex situ XRD, operando XRD, electrochemical cycling.</p> <p>Abstract: Sodium-ion-batteries (SIB) are a low-cost alternative to currently used lithium-ion batteries (LIB) but suffer from poor cycling stability. Spray drying provides porous, hierarchically structured particles of cathode active material (CAM) in large amounts, suitable for up-scaling. Changing the chemical composition of the Na0.6Al0.11–xNi0.22–yFex+yMn0.66O2 layered oxides under identical synthesis conditions leads to differences in particle morphology, conductivities, sodium vacancy ordering and phase transition, therefore influencing the electrochemical performance via several mechanisms. Here, a broad overview on these changes for samples with variable nickel and iron content is presented. With increasing iron content, the particle porosity is reduced and lower of initial capacity is received for most cycling windows. Substituting half of the original Ni amount with Fe still leads to high capacities and improved cycling stability. The influence of Al as electrochemical inactive element becomes visible in stabilised cycling stability as well.</p>
The structure of simple satellite variation in the human genome and its correlation with centromere ancestry (Supplemental Data)
<p>Accompanying <a href="https://github.com/is-the-biologist/1KGP_SATS" target="_blank" rel="noopener">Github</a></p> <p><strong>Supplemental File 1.</strong> BLAST results of k-mer concatemers against T2T-CHM13-v2.0.</p> <p><strong>Supplemental File 2.</strong> Annotations of centromeres, and telomeres of T2T-CHM13-v2.20. Table of abundance of k-mers in annotated regions as numpy file from BLAST hits. Abundance of k-mers across genome in 100kb bins from BLAST hits as .npz files accessible by example:</p> <p> import numpy as np<br> dense = np.load("filename.npz")<br> dense["chr1"]<br> <br><strong>Supplemental File 3</strong>. Table of pairwise R2 between simple satellites and table of pairwise interspersion OR between simple satellites. Folder containing QQ plots of negative binomial fit of satellite copy number distribution used to qualitatively asses model fit.</p> <p><strong>Supplemental File 4. </strong>Materials and results of cenGRM analysis. Boundaries used for centromeric regions of each cenGRM, cenGRMs in GCTA format, and tables with the results of cenGRM GCTA runs. Also provide pdfs of the dendrograms/heatmaps produced from UPGMA clustering of each cenGRM. </p> <p><strong>Supplemental File 5</strong> Non-human significant BLAST hits from BLAST-ing k-mer concatamers to non-human sequences.</p> <p><strong>Supplemental Table 1.</strong> Copy number normalized to 1x depth given GC bias of 126 most abundant satellites analyzed in paper in each individual. Additional columns represent metadata of the individual:</p> <ul> <li>instrument: sequencer instrument name used to sequence library.</li> <li>run: sequencer run of the library.</li> <li>flow: flowcell ID of the ibrary.</li> <li>pop: 1,000 Genomes Project population ID.</li> <li>superpop: 1,000 Genomes Project superpopulation ID.</li> <li>reads: average autosomal read depth of the library.</li> </ul> <p><strong>Supplemental Table 2. </strong>Copy number normalized to 1x depth given GC bias of the top 126 most abundant satellites analyzed in paper in each individual of the 1KGP, plus estimates of the same satellites in CHM13 short-read libraries subsampled from 18x-0.5x, 18x depth simulated library of the T2T-CHM13v2.0 assembly analyzed using k-Seek, and Tandem Repeat Finder results of the T2T-CHM13v2.0 asembly <a href="https://doi.org/10.1126/science.abk3112" target="_blank" rel="noopener">Hoyt 2022</a>.</p> <p><strong>Supplemental Table 3.</strong> Copy number normalized to 1x depth given GC bias of all tandem repeats with k-mer <= 20 (6,309) found collectively in the CHM13 short-read libraries subsampled from 18x-0.5x, 18x depth simulated library of the T2T-CHM13v2.0 assembly analyzed using k-Seek, and Tandem Repeat Finder results of the T2T-CHM13v2.0 asembly <a href="https://doi.org/10.1126/science.abk3112" target="_blank" rel="noopener">Hoyt 2022</a>.</p>
Data from: emergence of structure in plant-pollinator networks: low floral resource constrains network specialisation
<p>Specialisation enhances the efficiency of plant-pollinator networks through the exchange of conspecific pollen transfer for floral resources. Floral resources form the currency of plant-pollinator interactions, but the understanding of how floral resources affect the structure of plant-pollinator networks remains modest. Previous theory predicts that optimally foraging animal species will specialise to improve resource acquisition under high resource availability. Although floral resource availability depends on both the plant production and animal consumption of the resources, previous work has assumed that production and availability to be equivalent. This potentially may have led to erroneous inferences on the effect of resource availability on specialisation. We develop a mutualistic Lotka-Volterra consumer-resource model to investigate the influence of floral resource availability on plant-pollinator network structure. The model incorporates animal adaptive foraging behaviour, floral resource dynamics, and density-dependent dynamics. Specialisation, nestedness and modularity of simulated networks generated from the model under a wide range of parameters were explained using the Generalised Linear Model. We found that the distinction between floral resource dynamics and plant density dynamics was necessary for partial specialisation of plant-pollinator networks. This is because floral resource dynamics constraint animal preference due to its depletion by animal species. Floral resource abundance had a positive effect on network specialisation, but animal density had a negative effect on network specialisation. Floral resource dynamics thus play key roles on the structure of plant-pollinator network, distinctive from plant species density dynamics.</p>
Figs 53–56. Metaventrite structure. 53 in ON SPLITTING OF THE GENUS NOTOCUPES (COLEOPTERA: ARCHOSTEMATA): NEW DATA ON MORPHOLOGY AND TAXONOMY
Figs 53–56. Metaventrite structure. 53 – Conexicoxa crassa; 54 – Notocupes excellens; 55 – Rhabdocupes oxypygus; 56 – Rhabdocupes issykkulensis. Red indicates paracoxal suture and posterior margin of metaventrite. Scale bar = 1 mm.
Computed data for "Structural and Electronic Impacts of the Axial Substitution at the Phosphorus Center of C(sp3)-Bridged P-Heterotriangulenes"
<p>Computed structures and TD-DFT raw data of the article "Structural and Electronic Impacts of the Axial Substitution at the Phosphorus Center of C(sp3)-Bridged P-Heterotriangulenes" published in Eur. J. Org.Chem. <a href="https://doi.org/10.1002/ejoc.202400368">https://doi.org/10.1002/ejoc.202400368</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.