Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
311
datasets available to search
ShareScore release 0.9.0
Dataset results
311 results for “new dataset”
FIGURE 5 in A new morphological dataset reveals a novel relationship for the adzebills of New Zealand (Aptornis) and provides a foundation for total evidence neoavian phylogenetics
FIGURE 5. Synapomorphies for the pelvis of Aptornis defossor (AMNH 7300, A) and Psophia obscura (AMNH 2671, B). The pelvises are shown in ventral view. Scale bars are different for each specimen and are shown below each specimen. Labels correspond to synapomorphies, with character numbers followed by character states in parentheses. Abbreviations: cio, crista iliaca obliqua; ili, preacetabular ilium; ish, postacetabular ischium; pil, postacetabular ilium; syn, synsacrum
FIGURE 1 in A new morphological dataset reveals a novel relationship for the adzebills of New Zealand (Aptornis) and provides a foundation for total evidence neoavian phylogenetics
FIGURE 1. Strict consensus cladogram of nine most parsimonious trees (length: 2038, CI: 0.2498, RI: 0.5337, RC: 0.1333, HI: 0.7502) from analysis of our new morphological dataset of 40 taxa and 368 characters in PAUP*. Results support optimization of an Aptornis defossor + Psophia obscura sister group. Synapomorphies are detailed in table 2. Extinct taxa are denoted with daggers. Bootstrap support values greater than 50% are annotated above branches, with branch length ranges reported below branches.
FIGURE 3. A in A new morphological dataset reveals a novel relationship for the adzebills of New Zealand (Aptornis) and provides a foundation for total evidence neoavian phylogenetics
FIGURE 3. A. Resulting tree from Bayesian analysis of 32 RAG1 and RAG2 sequences. Clade credibility values greater than 90% are annotated above branches. Core Gruiformes, Ralloidea, and Gruoidea are well supported with 100% clade credibility values. The scale bar at the bottom of the tree denotes branch length. The data were run in MrBayes for 2,000,000 generations.
FIGURE 3 in A new morphological dataset reveals a novel relationship for the adzebills of New Zealand (Aptornis) and provides a foundation for total evidence neoavian phylogenetics
FIGURE 3 (continued). B. Resulting tree from Bayesian analysis of our new dataset of 40 taxa and 368 osteological characters combined with 32 sequences of RAG1 and RAG2 nuclear genes. Clade credibility values greater than 90% are annotated above branches. Extinct taxa are denoted with daggers. The scale bar at the bottom of the tree denotes branch length. The data were run in MrBayes for 1,100,000 generations.
A new sampling capability for uncertainty quantification in the Ice-sheet and Sea-level System Model v4.19 using Gaussian Markov random fields -- Datasets and results
<p>Data archives for test experiments (Section 3) and Pine Island Glacier application (Section 4) from the manuscript "Kevin Bulthuis and Eric Larour, A new sampling capability for uncertainty quantification in the Ice-sheet and Sea-Level System Model v4.19 using Gaussian Markov random fields"</p> <p>Source code is available at https://doi.org/10.5281/zenodo.5532775.</p>
Ngaruroro River, New Zealand - Geomorphic Change Detection - Example Dataset
<p>A simple<a href="https://gcd.riverscapes.net/Tutorials/example-data-sets.html"> Example GCD Dataset </a>illustrating topographic change detection on a long (31 km) dataset. Great for learning about analyses with Directional Masks in GCD.</p> <p>Dataset is from:</p> <ul> <li>31km river on the <a href="https://www.google.com/maps/place/39%C2%B035'58.6%22S+176%C2%B043'23.7%22E/@-39.6060374,176.6490462,27291m/data=!3m1!1e3!4m5!3m4!1s0x0:0x0!8m2!3d-39.599602!4d176.723239">north island of New Zealand</a></li> <li>Two LiDAR surveys</li> <li>2m cell resolution</li> </ul> <p>Dataset includes raw data to run exercises, as well as full *.gcd projects that can be opened. </p>
A New Reactivity Control Approach for Circulating Fuel Reactors - Dataset
<p>Dataset associated to the conference paper "A New Reactivity Control Approach for Circulating Fuel Reactors", Proceedings of the International Conference Nuclear Energy for New Europe (NENE2021), Bled, Slovenia, September 6–9, 2021</p>
Dataset: A New Approach to Scribal Abbreviation in the Bestiaire in Merton College Library, MS 249
<p>An accompanying dataset for analysis of scribal abbreviation in Oxford, Merton College Library, MS 249, ff. 1<sup>r</sup>-10<sup>v</sup>.</p>
Dataset for: The redlegged earth mite draft genome provides new insights into pesticide resistance evolution and demography in its invasive Australian range
<p>Data and analyses for Thia et al. "The redlegged earth mite draft genome provides new insights into pesticide resistance evolution and demography in its invasive Australian range" submitted to <em>Journal of Evolutionary Biology</em>.</p> <p>This repository comprises data and scripts used to replicate the analyses in this paper.</p> <p>The goals of this study were to: (1) assemble a draft reference genome for <em>Halotydeus destructor</em>; (2) perform a comparative analysis of acetylcholinesterase genes among different agricultural arthropod pests; (3) characterise the population genetic patterns among Australian <em>H. destructor</em> populations; and (4) perform demographic modelling to understand the evolutionary relationships between eastern and western populations of <em>H. destructor</em> in Australia.</p>
Dataset describing the vulnerability of the Nigerian Power system via a new voltage stability pointer
<p>The Nigerian power network (NGP), 28-bus 330 kV dynamic stability assessment using a new voltage stability pointer (NVSP). Different case studies were considered and presented on dynamic stability including the contingency analysis.</p> <p> </p>
Dataset of the Article "Reconstruction of the unbinding pathways of new inhibitors of the SARS-CoV-2 Papain-like protease using molecular dynamics simulation"
<p>This dataset contains concatenated trajectory files of the SuMD simulation of the unbinding pathways of the new inhibitors for SARS-CoV-2 papain-like protease. This data will be published in an article titled: "<strong>Reconstruction of the unbinding pathways of new inhibitors of the SARS-CoV-2 Papain-like protease using molecular dynamics simulation".</strong></p>
Datasets related to agroinoculation and transmission experiments with Tomato Leaf Curl New Delhi Virus Spain Strain between zucchini cv. Milenio and Ecballium elaterium plants by Trialeurodes vaporariorum and Bemisia tabaci MED whiteflies
<p>These datasets support the data provided directly in the scientific publication titled "Tomato Leaf Curl New Delhi Virus Spain Strain Is Not Transmitted by Trialeurodes vaporariorum and Is Inefficiently Transmitted by Bemisia tabaci Mediterranean between Zucchini and the Wild Cucurbit Ecballium elaterium", published in the <a href="https://www.mdpi.com/2075-4450/14/4/384">Insects journal</a>. </p> <p>The figure included in the datasets highlights agroinoculation experiments with tomato leaf curl New Delhi virus Spain strain using different source plant–target plant–whitefly species combinations. ToLCNDV (DNA-A and DNA-B) was detected by molecular hybridization after tissue printing on nylon membranes using specific digoxigenin-labelled DNA probes for each viral component. (<strong>A</strong>) Prints of zucchini plants inoculated using Trialeurodes vaporariorum (T. v.) or Bemisia tabaci MED (B. t.) as a vector and infected zucchini plants as the virus source. (<strong>B</strong>) Prints of Ecballium elaterium plants using B. t. as a vector and infected zucchini plants as the virus source. (<strong>C</strong>) Prints of zucchini plants using B. t. as a vector and infected E. elaterium plants as the virus source. Blots from one of the two independent experiments (Experiment 1 in the Table) carried out for each source plant–target plant–whitefly species combination are shown. Prints of ToLCNDV-infected zucchini plants included as positive controls are outlined with red squares.</p> <p>The table included in the datasets show the results of transmission experiments of tomato leaf curl New Delhi virus Spain strain between zucchini cv. Milenio and Ecballium elaterium plants by Trialeurodes vaporariorum and Bemisia tabaci MED. Infection was determined by tissue-printing and molecular hybridization with digoxigenin-labelled ToLCNDV DNA-A and DNA-B probes. Blots from experiment (Exp.) 1 of each source plant–target plant–whitefly species combination are shown in the Figure.</p> <p>The provided datasets are further discussed and interpreted in detail, as well as their subsequent results, in the scientific publication.</p> <p>This research was conducted within the VIRTIGATION project, which is part of the EU Open Research Data pilot. This project has received funding from the European Union’s Horizon 2020 research and innovation program under grant agreement No. 101000570.</p>
Dataset to reproduce firgures for the paper "A unifying method to study Respiratory Sinus Arrhythmia dynamics implemented in a new toolbox"
<p>Dataset provided to reproduce figures <br> for the paper "A unifying method to study Respiratory Sinus Arrhythmia dynamics implemented in a new toolbox"</p> <p>Jupyter notebooks are available here:<br> https://github.com/samuelgarcia/physio_benchmark</p> <p>Human dataset<br> =============</p> <p>Context: A research aimed to decipher the impact of respiration on brain oscillations</p> <p>Data collection methods: ECG and Respiration of 15 healthy adults subjects <br> (age : 30.9 +/- 9.5 yo).All participants gave informed consent to take part to the study, and all experiments<br> were approved by the national french committee (CPP number 4090). They were sitting quietly and instructed just<br> to relax. Recording lasted 5 minutes. Respiration signal was recorded from a nasal sensor<br> (Sensortechnics GmbH, Puchheim , Germany) at a sampling rate of 1000 Hz, amplified by actiCHamp<br> Plus amplifier (Brain Products GmbH, Gilching, Germany). ECG signal was recorded from 3 skin electrodes<br> (right forearm, left forearm, left iliac region), at a sampling rate of 1000 Hz (same amplifier).</p> <p>Structure of files: tabular separated values text files.<br> The first columns correspond to the ECG signal, the second is the respiratorysignal.<br> The sampling rate is 1000Hz</p> <p>Data manipulations: The original dataset has longer durationand and contain channels (EEG).<br> This sub-dataset was extracted from the original using the neo python package from the VHDR brain product format.<br> Signal tarces haven't been preprocessed they correspond to the "raw" signal.</p> <p><br> Data confidentiality and permissions: Experiments were approved by the national french committee (CPP number 4090)</p> <p><br> Animal dataset<br> ==============</p> <p>Context: The dataset was recorded to validate the device telemetric jacket from Etisense.</p> <p>Data collection methods: ECG and Respiration of 1 adult rat were recorded. Recording lasted 30 seconds during<br> freely behaving. Respiration signal and ECG were recorded from a thoraco-abdominal telemetric<br> jacket at which it was habituated before. Recorded were done at a sampling rate of 500 Hz, amplified by<br> Etisense acquisition unit (Etisense, MedTech company, Lyon, France).</p> <p>Structure of files: tabular separated values text files.<br> The first columns correspond to the ECG signal, the second is the respirator signal.<br> The sampling rate is 500Hz</p> <p><br> Data manipulations:<br> The dataset was extracted from the original HDF5 structure.<br> The ECG signal correcpond to the "raw" traces from the HDF5 files.<br> The respiratory signal was originaly sample at 200Hz on the device and resample with linear interpolation<br> to 500Hz to be easy aligned with the ECG signal.</p> <p>Data confidentiality and permissions: Experiments were carried according to the ethical guidelines of the<br> European Communities Council Directive of 24 November 1986 (86/609/EEC), as well as the approval 16979 of the<br> Lyon 1 University CEEA-55 ethical committee and of the Ministry of Higher Education, Research and Innovation.<br> </p>
Supplementary data and scripts for "New insights into the relationship between mass eruption rate and volcanic column height based on the IVESPA dataset"
<p>Supplementary tables and MATLAB scripts associated with the manuscript "New insights into the relationship between mass eruption rate and volcanic column height based on the IVESPA dataset".</p>
Seismic dataset for Ruapehu and Whakaari volcanoes in New Zealand
<p>RSAM, MF, HF and DSAR time series for Ruapehu stations FWVZ over the 14 years explored, and for Whakaari stations WIZ over 9 years.</p> <p>Computing datastreams: we harnessed seismic data from a vertical component station for each individual volcano. We applied data processing techniques that resulted in the generation of four distinct time series, with a sampling interval of 10 minutes. Various measures were employed to capture different aspects of the seismic signal. The first measure, known as the Real-time Seismic Amplitude Measurement (RSAM), was obtained by calculating the 10-minute moving average of the velocity recorded by the vertical station This signal was then subjected to bandpass filtering within the frequency range of 2 to 5 Hz, which focuses on tremor signal of frequent volcanic origin while excluding ocean noise at lower frequencies. Similarly, the Median Frequency (MF) and High Frequency (HF) measures were derived using a comparable approach to RSAM, but with specific bandpass filtering applied. MF was obtained by filtering the signal within the frequency range of 4.5 to 8 Hz, while HF was obtained by filtering within the frequency range of 8 to 16 Hz. The 4.5 Hz threshold between RSAM and MF reflects an assumption that tremor mostly radiates energy below 4.5 Hz. To exclude this effect and explore attenuation related to permeability change (such as sealing), this frequency value is used as a threshold. Lastly, the Displacement Seismic Amplitude Ratio (DSAR) was calculated as the ratio of the integrals of the MF and HF signals. High values of DSAR have been inferred to correlate with high gas levels in the edifice, suggesting either reduced fluid motion and/or trapping that has led to a gas-accumulation.</p>
Supplementary dataset for "Evaluating Graph Neural Networks for Link Prediction: Current Pitfalls and New Benchmarking"
<p>The supplementary dataset for the paper "Evaluating Graph Neural Networks for Link Prediction: Current Pitfalls and New Benchmarking". We include the splits for cora, citeseer, and pubmed, the hard negative samples, and the node2vec embeddings. We also include a jupyter file <em>read_data.ipynb</em> to show how to read the non-txt file.</p> <ul> <li>heart_test_samples.npy, heart_valid_samples.npy: the heard negative samples</li> <li>*-n2v-embedding.pt: node2vec embeddings</li> <li>test_samples_index.pt, valid_samples_index.pt: the node index of the selected samples in ogbl-ppa under HeaRT</li> <li>gnn_feature: the input feature of cora, citeseer, pubmed</li> </ul> <p>More details for our code and how to use the dataset are on the code repository: https://github.com/Juanhui28/HeaRT .</p>
A new version of ManyTypes4Py Dataset (Version 0.8)
<p>A new version of the <em>ManyTypes4Py</em> dataset could be used to train and evaluate the TypePy model</p> <p>For the scripts to make and reproduce it please refer to <a href="https://github.com/LangFeng0912/build_MTV0.8">Github Repo</a></p> <p> </p>
Dataset: Paranita, a new genus of spiders from northeastern Argentina (Araneae, Trachelidae)
<p>DNA alignments, morphological data and phylogenetic tree. Alignments of 6 DNA markers and morphological data for the phylogenetic analysis of teh genus Paranita. For the phylogenetic analysis, we composed a dataset combining eight traditional target markers from the analysis of Wheeler et al. (2017) (12s, 16s, 18s, 28s, co1, H3), plus sequences from other sources and new co1 sequences for P. paulae (BOLD CORAR075, GenBank OR515542, MACN-Ar 30271) and Trachelopachys sericeus (BOLD SPDAR1264-15, GenBank OR515543, MACN-Ar 34546). Some markers were retrieved as bycatch from Sequence Read Archive (SRA) phylogenomic sequences (see publication for details). We used the morphological data accumulated in the datasets of Azevedo et al. (2022a) and Ramírez (2014). The analyses under maximum likelihood were made with IQ-TREE 2.2.0 (Minh et al. 2020). Models for each target-gene were selected by Bayesian information criterion with ModelFinder (Kalyaanamoorthy et al. 2017). The models selected for the sequence data were as follows: TIM2+F+G4 (12s and 16s), TNe+R2 (18s, co1-2, h3-1, h3-2), GTR+F+I+G4 (28s), GTR+F+I+G4 (co1-1), GTR+F+I+G4 (co1-3), GTR+F+I+G4 (h3-3). The morphological data was partitioned into two datasets, one with the unordered characters, another with the ordered ones, and analyzed with the Mk and Mk-ordered models, respectively, both with correction for ascertainment bias for the absence of invariant characters. Prior to analysis, all invariant characters were removed, and polymorphic entries were replaced by missing entries. The branch support was estimated with 1000 rounds of ultrafast bootstrap (Hoang et al., 2018). The analyses under maximum parsimony were made with TNT v 1.6 (Goloboff & Catalano, 2016) under equal weights using an exact search of implicit enumeration. Branch support was measured with 1000 rounds of jackknifing, representing frequencies over the optimal tree. See publication for references and details.</p>
NEMARCO project: Dataset for the publication "Development of a new manufacturing route for NiCrSiFeB alloys by Direct Energy Deposition Laser Beam process (LMD)"
<p><strong>LMD dataset</strong></p> <p>This dataset gathers data from different parts of the Laser Metal Deposition metal Additive Manufacturing process (DED-LB). The dataset covers not only the process development data for samples manufacturing and monitored data of the melt pool size during the process, but also the metrics associated to the powder feedstock consumption, energy consumption and process efficiency.</p> <p><strong>Motivation</strong></p> <p>Nickel-based NiCrSiFeB alloy (Ni-Cr-Si-B self-fluxing family) are excellent candidates for replacing Cobalt-based alloys in aeronautical components such as sealing rings, valve seats, sliding bearing seats, etc. In this type of components, commonly manufactured by centrifugal casting and conventional processes, high temperature wear and stiffness under complex thermo-mechanical stresses cause lack of sealing and an increase in the wear rate. Metal additive manufacturing by direct laser metal deposition with powder (p-LMD) is presented as a potential manufacturing route for the complex processing of this type of alloys. This research work deals with the development of a new manufacturing route using p-LMD that ranges from the proper selection of the chemical composition for the starting powders, the development of the LMD process parameters to tackle the challenges associated to the wide solidification range and crack susceptibility of Ni-Cr-Si-B alloys, its monitoring and control, as well as the post- processing required to achieve the manufacture of aeronautical components.</p>
Dataset for the article "Beyond PLFA: Concurrent extraction of neutral and glycolipid fatty acids provides new insights into soil microbial communities"
<p>The following are data and code used for statistical analysis and figure plotting in the manuscript</p> <p>Gorka et al. (2023) "Beyond PLFA: Concurrent extraction of neutral and glycolipid fatty acids provides new insights into soil microbial communities", Soil Biology and Biochemistry</p> <p>It contains the following files:</p> <p>1. Pure lipid standard data</p> <ul> <li>Total ion chromatogram (TIC) area data (<strong>area.csv</strong>)</li> <li>Assignment of lipids that the measured fatty acids originate from (<strong>LipidClass.csv</strong>)</li> <li>An R script reproducing the calculations and plotting for Fig. 2 and Fig. S1 (<strong>pure_lipids.R</strong>)</li> </ul> <p>2. Microbial pure culture fatty acid data data</p> <ul> <li>TIC area data of the PLFA, NLFA, and GLFA data from pure culture extracts (<strong>area.csv</strong>)</li> <li>Files needed for calculating the data and assigning taxonomic groups in the R code (<strong>weights.csv</strong>, <strong>C_atoms.csv</strong>, <strong>species_list.csv</strong>)</li> <li>An R script reproducing the calculations and plotting for Fig. 3, Fig. 4, Fig. S2, and Fig. S3 (<strong>pure_cultures.R</strong>)</li> </ul> <p>3. Soil fatty acid data</p> <ul> <li>Absolute abundance data in nmol C g<sup>-1</sup> dry weight of the PLFA, NLFA, and GLFA data from soil extracts (<strong>nmolC.csv</strong>)</li> <li>Taxonomic group assignments of fatty acids needed to run the R code (<strong>phylum.csv</strong>)</li> <li>An R script reproducing the calculations and plotting for Fig. 5, and Fig. S4 (<strong>soil.R</strong>)</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.