Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Data Set "Benchmarking of Vibrational Exciton Models Against Quantum-Chemical Localized-Mode Calculations"
<p>This data set accompanies the publication "Benchmarking of Vibrational Exciton Models Against Quantum-Chemical Localized-Mode Calculations" <br>by Anna M. van Bodegraven, Kevin Focke, Mario Wolter, and Christoph R. Jacob <br>(TU Braunschweig, Germany) </p> <p>It contains the following files:</p> <p><br>Directory '01_AIM':</p> <p> - input (structure.pdb, topol.top and *_input.txt) and results (*.log and<br> Hamiltonian/AtomPos/Dipole/Parameters.txt) from frequency calculations<br> with the Amide-I-maps (AIM) program for six polypeptide test cases<br> (1gpB_310, ala_310, ala_hairpin, ala_helix, ala_strand, and trpzip), <br> each in vacuo or water with three different maps (Jansen, Skinner, <br> Tokmakoff) for 11 snapshots based on a MD run.<br> To rerun the calculations, you will have to change the paths in <br> the *_input.txt files (topfile, trjfile, sourcedir) accordingly<br> <br>Directory '02_SNF':</p> <p> - results (*.dat, *.out, coord and control) from frequency calculations <br> using Turbomole, SNF, and LocVib for six polypeptide test cases<br> (1gpB_310, ala_310, ala_hairpin, ala_helix, ala_strand, and trpzip) <br> each in vacuo or water for snapshots based on an MD run.<br> <br>Directory '03_NMA':</p> <p> - coordinates and results from pyADF for NMA molecules each alligned <br> with a peptide bond from the six polypeptide test cases<br> (1gpB_310, ala_310, ala_hairpin, ala_helix, ala_strand, and trpzip) <br> each in water for 11 snapshots based on a MD run and input <br> (*_input.txt) and results (*.log and Hamiltonian/AtomPos/Dipole/Parameters.txt) <br> from calculations with the Amide-I-maps (AIM) program<br> <br>Directory '04_Handling_Data':</p> <p> - contains all used notebooks to extract the data, plot the figures <br> and calculate the errors<br> - RMSD_Error_vacuo/water.ipynb is used to calculate the overall shifts <br> and generates a map-dependent mean value to shift the frequencies of AIM<br> - Frequencies.ipynb and Couplings.ipynb are used to plot the figures <br> - RMSD_values_vacuo/water.ipynb show the calculations for the statistical <br> analysis<br> To use the notebooks, start with the Dictionary_setup_for_data_for_paper.ipynb <br> to set up the main dictionary from the calculated data</p>
Data sets used in "Is there a tropical response to recent observed Southern Ocean cooling?"
The dataset contains processed data for the GRL publication (to be submitted) "Is there a tropical response to recent observed Southern Ocean cooling?" The dataset includes global maps of linear trends of various fields presented in the figures, as well as pattern correlation coefficients for the three experiment ensembles using CESM1: Large Ensemble, Tropical Pacific Pacemaker Experiment, and Southern Ocean Pacemaker Experiment.
Merged Hadley-OI sea surface temperature and sea ice concentration data set
<p>The merged Hadley-OI sea surface temperature (SST) and sea ice concentration (SIC) data sets were specifically developed as surface forcing data sets for AMIP style uncoupled simulations of the Community Atmosphere Model (CAM). The Hadley Centre's SST/SIC version 1.1 (HADISST1), which is derived gridded, bias-adjusted in situ observations, were merged with the NOAA-Optimal Interpolation (version 2; OI.v2) analyses. The HADISST1 spanned 1870 onward but the OI.v2, which started in November 1981, better resolved features such as the Gulf Stream and Kuroshio Current which are important components of the climate system. Since the two data sets used different development methods, anomalies from a base period were used to create a more homogeneous record. Also, additional adjustments were made to the SIC data set.</p>
Data sets for Shiburah et al., "The absence of the Leishmania major telomerase TERT component links telomeres and cell homeostasis with infectivity"
<p>These files correspond to the figures and information contained in Shiburah et al., "The absence of the <em>Leishmania major</em> telomerase TERT component links telomeres and cell homeostasis with infectivity"</p>
Structures of FDA-approved drugs and their active metabolites and data sets of experimental PD and PK properties
<p>Data sets are extracted from the 2024 release of the e-Drug3D Database (2118 FDA-approved drug structures)</p> <ol> <li><strong>e-Drug3D_2118.zip </strong>(contains e-Drug3D_2118.sdf)<strong> -</strong> <strong>Chemical Structures</strong> - The e-Drug3D collection in SDF format file - one 3D conformer; ionization of carboxylic acid, phosphate, phosphonate, phosphonoamide, amidinium and guanidinium groups. The datablock contains the ID, name (INN), CAS number and Status.</li> <li><strong>e-Drug3D_2118_PK.csv - </strong><strong>Pharmacokinetics</strong> - Column/field value is separated by a semicolon. It contains the e-Drug3D ID, INN (drug name), CAS number, year of approval, Status, is_or_has a metabolite, routes of administration, Volume of distribution (VD), Clearance (Cl), Plasma Protein Binding (PPB), Half-life (t1/2), Bioavailability (F), Cmax/Tmax, comment on solubility.</li> <li><strong>e-Drug3D_2118_PD.csv -</strong> <strong>Pharmacodynamics</strong> - Column/field value is separated by a semicolon. It contains the e-Drug3D ID, INN (drug name), CAS number, year of approval, Status, Primary target, ATC code(s), PDB codes and main list of drug targets.</li> <li><strong>e-Drug3D_2118_RD.csv -</strong> <strong>FDA Registration Data</strong> - Column/field value is separated by a semicolon. It contains the ID, name (INN), CAS number, First year of approval, Status, <a href="http://www.knapsackfamily.com/knapsack_core/top.php">KNApSAcK</a> or <a href="https://www.npatlas.org">NPAtlas</a> Id if natural product, all associated NDA numbers [FDA approval number, name of the label file in PDF format, company name, year of approval and commercial name of the drug] and the Indication/Therapeutic class information.</li> <li><strong>labels.tar.gz</strong> - The drug label files in PDF format (compressed directory). A label file is named with the NDA number. The NDA number is the approval number assigned by the FDA. A drug may possess several NDA numbers (see the above e-Drug3D-RD data set).</li> </ol>
[VERSION 2] Data set and analytic codes supporting "How do management decisions impact butterfly assemblages in smallholding oil palm plantations in Peninsular Malaysia?"
<p>This is <strong>VERSION 2</strong> of data set and analytic codes (with a meta data [see the meta data from VERSION 1]) supporting "How do management decisions impact butterfly assemblages in smallholding oil palm plantations in Peninsular Malaysia?". We investigated the impacts of replanting and alternative replanting decisions (replanting with monoculture versus polyculture oil palm plantations) on within-plantation environmental conditions and butterfly assemblages (diversity, density, and composition). We also assessed the effects of habitat structure and complexity within plantations on butterfly assemblages. Apart from "BantingButterflies_ButterflyData", other data are the same as in VERSION 1.</p><p><strong>## List of changes:</strong></p><p># 1. <i>Tirumala septentrionis </i>was not included in the analyses because it should have been <i>Ideopsis vulgaris</i> (had been corrected),</p><p># 2. PC5 and PC6 (from PCA) were considered as predictors for the GLMs,</p><p># 3. The Mantel test was added.</p><p><strong>## Other notes:</strong></p><p># 1. Older version of ggiNEXT could work with facet.var = "site", now it needs to be "Assemblage"</p><p># 2. Older version of ggiNEXT could work with facet.var = "order", now it needs to be "Order.q"</p><p># 3. "set.seed(42)" function was used before running "iNEXT", ANOSIM, and the Mantel test to get reproducible outputs (exactly the same outputs every time each function is run).</p><p><strong>Funding and research permission:</strong> Jardine Foundation, the Cambridge Trust, and Tim Whitmore Fund provided funding for MFH, the Biotechnology and Biological Sciences Research Council (BBSRC) funded JS (USN: 304338625), and BBSRC (BB/T012366/1) provided funding for the establishment of the plots and surveys of environmental parameters. Research permission was provided by the Economic Planning Unit (EPU) of Malaysia's Prime Minister's Department for MFH (Ref: EPU 40/200/19/3727) and JS (Ref: MEA 40/200/19/3705).</p>
Numerical data set belonging to: 'A Finite Volume Parallel Adaptive Mesh Refinement Method for Solid-Liquid Phase'
<p>This data set corresponds to the paper 'A Finite Volume Parallel Adaptive Mesh Refinement Method for Solid-Liquid Phase Change', submitted to Numerical Heat Transfer, Part A: Applications. The numerical data is included in VTK format (to be read by paraView) for the following cases:</p><p>1) 2D Gallium melting in a rectangular cavity (70x50 elements, 140x100 elements, 280x200 elements, 560x400 elements, 1120x800 elements and adaptive mesh)</p><p>2) 3D Gallium melting in a hexagonal cavity (adaptive mesh)</p><p>3) 2D freeze-plug (both steady-state and melting transient): 110x300 elements, 220x600 elements, 440x1200 elements and adaptive mesh)</p><p>Due to the size of the data-set, the data has been split over 8 tar archives featuring a gzip compression. To unpack the data, run the command: cat paper_<i>amr</i>_<i>data.</i>tar.gz.* | tar xzvf -</p><p> </p>
Data set from the publication DOI: 10.1200/JCO.22.01748
<p>This is the original dataset including anatomised patient codes, overall survival, progression-free survival, PD-L1 TPS score, PD-1/PD-L1 interaction state (measured by QF-Pro technology of HAWK Biosystems, mean, median and upper quartile values given), patient status (dead or alive).</p>
Direct observation of chirality-induced spin selectivity in electron donor–acceptor molecules. Open data set
<p>Data supporting the original figures 2 and 4 of the related publication.</p>
Data set from Ediacaran Mirassol d´Oeste Formation
<p>These data are from geochemical analysis from the Paleoenvironmental evolution of the Ediacaran Mirassol d´Oeste Formation, Araras-Alto Paraguai Basin, central part of Brazil. These results are submitted to Global and Planetary Change.</p>
Data set for "Cyclophospholipids enable a protocellular life cycle"
<p>Publication in ACS Nano can be found <a href="https://doi.org/10.1021/acsnano.3c07706">here</a>.</p><p>Toparlaka ÖD, Sebastianelli L, Egas Ortunoc V, Karkic M, Szostak JW, Krishnamurthy R, Mansy SS (2023) Cyclophospholipids enable a protocellular life cycle. ACS Nano 17, 23772–23783. DOI: 10.1021/acsnano.3c07706</p>
Automated bio-AFM generation of large mechanome data set and their analysis by machine learning to classify prostatic cell lines_Training base 100 PC3-GFP
Open the record for dataset details and reuse information.
EOLSHIPS: A Global Data Set of End-of-Life Ships, 2017-2022
Open the record for dataset details and reuse information.
Data underpinning "Chaotic fluctuations in a universal set of transmon qubit gates"
<p>Transmon qubits arise from the quantization of nonlinear resonators, systems that are prone to the buildup of strong, possibly even chaotic, fluctuations. One may wonder to what extent fast gate operations, which involve the transient population of states outside the computational subspace, can be affected by such instabilities. We<br>here consider the eigenphases and -states of the time evolution operators describing a universal gate set, and analyze them by methodology otherwise applied in the context of many-body physics. Specifically, we discuss their spectral statistic, the distribution of time dependent level curvatures, and state occupations in- and outside the computational subspace. We observe that fast entangling gates, operating at speeds close to the so-called quantum speed limit, contain transient regimes where the dynamics indeed becomes partially chaotic. We find that for these gates even small variations of Hamiltonian or control parameters lead to large gate errors and speculate on the consequences for the practical implementation of quantum control.</p>
Experimental raw data sets associated with certified reference material BAM-P116 (titanium dioxide) for comparison of nitrogen and argon sorption, available in the universal adsorption information format (AIF)
<p>These data sets serve as models for calculating the specific surface area (BET method) using gas sorption in accordance with ISO 9277.<br>The present measurements were carried out with nitrogen at 77 Kelvin and argon at 87 Kelvin.<br>It is recommended to use the following requirements for the molecular cross-sectional area:<br>Nitrogen: 0.1620 nm²<br>Argon: 0.1420 nm²</p> <p>Expected specific surface area for nitrogen (BET): 305 to 345 m²/g<br>Expected specific surface area for argon (BET): 300 to 310 m²/g</p> <p>Titanium dioxides certified with nitrogen sorption and additionally measured with argon for research purposes were used as sample material.<br>The resulting data sets are intended to serve as comparative data for own measurements and show the differences in sorption behaviour and evaluations between nitrogen and argon.<br>These data are stored in the universal AIF format (adsorption information format), which allows flexible use of the data.</p>
Data for: DNA metabarcoding uncovers dispersal-constrained arthropods in a highly fragmented restoration setting
<p>Degraded areas are often restored through active revegetation, however recolonisation by animals is rarely engineered. Recolonisation may be rapid for species with strong dispersal abilities. However, poor dispersers, such as many flightless arthropods, may struggle to recolonise newly restored sites. Actively reintroducing or 'rewilding' arthropods may therefore be necessary to facilitate recolonisation and restoration of arthropod communities and the ecological functions they perform. However, active interventions are rare. The purpose of this study was twofold. First, we asked whether potential source remnant arthropod communities were dispersal-constrained and struggling to recolonise restoration sites. Second, we tested whether reintroducing entire arthropod communities from remnant populations would help dispersal-constrained species establish during farmland ecological restoration in southern Australia. Rewilding was conducted in summer 2018 by transplanting leaf litter, soil, and entire communities contained within it from remnant source populations into geographically isolated restoration sites, which were paired with untreated controls (n = 6 remnant, rewilding transplant, and control sites). We collected leaf litter and extracted arthropod communities 19 months after the initial rewilding event, then sequenced mite, springtail, and insect communities using a metabarcoding approach. Within all groups, community similarity decreased with spatial distance between sites, suggesting significant dispersal barriers. However, only mite communities showed a strong response to rewilding, which was expressed as increased compositional similarity towards remnant sites and greater species richness relative to controls. Our results demonstrate that many arthropod species may struggle to recolonise geographically isolated restoration sites and that full community restoration requires active interventions via rewilding.</p>
Data set for the article: An artificial neural network approach to finding the key length of the Vigenere cipher
<p>Data supporting the work in the article: An artificial neural network approach to finding the key length of the Vigen\`{e}re cipher.</p>
Data Set: Radio Frequency Senser (RFS) example waveforms, altitudes, and density plot
<p>This data set contains data for the paper entitled "Radio Frequency Sensor: radio frequency lightning detection in geostationary orbit", submitted to <em>Radio Science </em>in December 2023.</p> <p>This data set contains three types of RFS data. 1) The first are time domain waveforms of three lightning events, in both the RFS high band (116 – 142 MHz) and the RFS low band (10-60 MHz), sampled at 155 MHz. The waveforms are right-hand circularly polarized waveforms. 2) The second type of data is altitudes and locations of trans-ionospheric pulse pairs (TIPPs) over time. The locations were determined by time coincidence with geolocated World Wide Lightning Location Network strokes. 3) The third type is RFS events per square kilometer per year in latitude and longitude. The RFS event were located by time correlation to Earth Networks Global Lightning Network lightning strokes.</p> <p><strong>Data set 1:</strong></p> <p>Consists of six ASCII files – 3 high band & 3 low band example RFS right-hand circularly polarized waveforms. Each ascii file contains a header with the RFS event time in UTC, the label of “RFS high band (77.5 – 155 MHz)” or “low band (0 – 77.5 MHz)”, and sample rate (155 MHz). Data following the header are time samples of electric field in uV/m sampled at 155 MHz.</p> <p>Filenames are:</p> <p>RFS_waveform_HighBand_20230607_010803.txt</p> <p>RFS_waveform_HighBand_20230607_015553.txt</p> <p>RFS_waveform_HighBand_20230607_034055.txt</p> <p>RFS_waveform_LowBand_20230607_010803.txt</p> <p>RFS_waveform_LowBand_20230607_015553.txt</p> <p>RFS_waveform_LowBand_20230607_034055.txt</p> <p> </p> <p><strong>Data set 2:</strong></p> <p>Filename = ‘RFS_TIPPs_20230607_0100-0500_UTC.txt’</p> <p>1 ASCII comma separated value (CSV) file. Columns are:</p> <p>1. UTC date yyyy/mm/dd</p> <p>2. UTC seconds of day</p> <p>3. WWLLN-determined latitude (degrees, wwlln_latitude in header)</p> <p>4. WWLLN-determined longitude (degrees, wwlln_longitude in header)</p> <p>5. TIPP-estimated height (km, height in header)</p> <p> </p> <p><strong>Data set 3:</strong></p> <p>Filename = ‘RFS_map.csv’</p> <p>1 CSV file of a 2-dimensional data set.</p> <p>1. Row 1, Longitude (degrees, in 0.25-degree steps)</p> <p>2. Column 1, Latitude (degrees, in 0.5-degree steps)</p> <p>3. 2-D grid in latitude and longitude: events per square kilometer per year</p> <p>Notes: Since the RFS coverage range goes across longitude = -180/180 degrees, longitudes go from 147.75 to 180, then start at -180 to -12.75. Latitude range goes from -58.5 to 68.5, as there were no detected RFS events outside these latitudes.</p> <p>The three examples given in data set 1 are those shown in LA-UR-23-32419, Figure 2. The TIPP data in data set 1 is shown in LA-UR-23-32419, Figure 4, and comprises data from 07 June 2023 between 01:00-05:00 UTC. Data set 3 contains RFS event rates per sq. km per year for data from 1 March 2022 – 1 March 2023, with the caveats described in LA-UR-23-32419.</p>
Data set of strain, differential phase, RS value, strain and PGV.
<p>The data stored here is the observations of strain, differential phase, RS value, strain and PGV in the paper of ‘<span>Potential of Earthquake Strong Motion Observation Utilizing a Linear </span><span>Estimation</span><span> Method for Phase Cycle Skipping in Distributed Acoustic Sensing</span>’, by Katakami et al.</p>
Data set: Al-Biruni Earth Radius Optimization with Deep Transfer Learning based Scene Image Classification on Remote Sensing Imagery
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.