Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
424
datasets available to search
ShareScore release 0.9.0
Dataset results
424 results for “Compilation”
Fig. 58. Pterygoplichthys disjunctivus, 355 in The non-native freshwater fishes of Singapore: an annotated compilation
Fig. 58. Pterygoplichthys disjunctivus, 355 mm SL, Woodlands Pond (dorsal, lateral, and ventral views).
Fig. 14. Barbonymus gonionotus, 140.0 in The non-native freshwater fishes of Singapore: an annotated compilation
Fig. 14. Barbonymus gonionotus, 140.0 mm SL, trade material (specimen with aberrant body scale pattern).
Fig. 12. Barbodes semifasciolatus, 42.2 in The non-native freshwater fishes of Singapore: an annotated compilation
Fig. 12. Barbodes semifasciolatus, 42.2 mm SL, wild type from Vietnam (top); ca. 30 mm SL, xanthic trade material (bottom).
Fig. 8. Scleropages formosus, 32.0 in The non-native freshwater fishes of Singapore: an annotated compilation
Fig. 8. Scleropages formosus, 32.0 mm SL juvenile (top), 490 mm SL adult (bottom), Upper Peirce Reservoir.
Fig. 26 in The non-native freshwater fishes of Singapore: an annotated compilation
Fig. 26. Esomus metallicus, ca. 40 mm SL, trade material (note the diagnostic character of long barbels).
Fig. 48. Gyrinocheilus aymonieri, 136.0 in The non-native freshwater fishes of Singapore: an annotated compilation
Fig. 48. Gyrinocheilus aymonieri, 136.0 mm SL, trade material (top); 117.4 mm SL, xanthic variety, Bedok Reservoir (bottom).
Text-fig. 10. Phylogenetic relationships of Miocene hyaenodonts (for definitions of character states see Table 2). The data matrix was compiled in MacClade 4.05 and run in PAUP 4.0b10 (Macintosh version). We chose Cimolestes magnus CLEMENS et RUSSELL, 1965, (additional data from Lillegraven 1969), as the outgroup. The unordered and unweighted analysis produced 16 trees. a: Majority-rule consensus. b: Strict consensus. Consistency index (CI): 0.5882; Homoplasy index (HI): 0.4118; Retention index (RI): 0.7742. in New Hyaenodonts (Ferae, Mammalia) From The Early Miocene Of Napak (Uganda), Koru (Kenya) And Grillental (Namibia)
Text-fig. 10. Phylogenetic relationships of Miocene hyaenodonts (for definitions of character states see Table 2). The data matrix was compiled in MacClade 4.05 and run in PAUP 4.0b10 (Macintosh version). We chose Cimolestes magnus CLEMENS et RUSSELL, 1965, (additional data from Lillegraven 1969), as the outgroup. The unordered and unweighted analysis produced 16 trees. a: Majority-rule consensus. b: Strict consensus. Consistency index (CI): 0.5882; Homoplasy index (HI): 0.4118; Retention index (RI): 0.7742.
Compilation of LCF and TMF Data for Single-Crystal and Directionally-Solidified Ni-base Superalloys Loaded in <001>
<p>This Excel file contains a compilation of low cycle fatigue (LCF) and thermomechanical fatigue (TMF) data for single-crystal (SX) and directionally-solidified (DS) Ni-base superalloys that are typically used for blades and vanes in gas turbines. Most of the data is generated by strain-controlled uniaxial fatigue tests. Both continuous cycling and cycles with dwells (hold times) at either the maximum or minimum of the cycle are included in the dataset.</p> <p>Several different alloys with similar mechanical behaviors are included in the dataset. The dataset only includes materials that have undergone a solution and age heat treatment typical of that applied to blade and vane components. All the fatigue data in this dataset is limited to uniaxial cyclic loading in the <001> direction for SX and in the longitudinal direction for DS alloys. The data in this dataset has been extracted from the public domain. The sources of the data are included in the Excel file.</p> <p>This data is used to train a Probabilistic Physics-guided Neural Network (PPgNN) that outputs a strain-life curve, including the mean response and 95% confidence intervals. The INPUTS and OUTPUTS used in the PPgNN are highlighted in the Excel file.</p>
Anechoic and IR Convolution-based Auralization Data Compilation Ensemble (AIRCADE)
<p><strong>AIRCADE</strong> is a data-compilation ensemble, primarily intended to serve as a resource for researchers in the field of dereverberation, particularly for data-driven approaches. It comprises <strong><a href="https://zenodo.org/record/1188976#.ZDhNTHbMJPY">speech and song samples</a></strong>, together with <strong><a href="https://zenodo.org/record/3371780#.ZDhOC3bMJPZ">acoustic guitar sounds</a></strong>, with original annotations pertinent to emotion recognition and Music Information Retrieval (MIR). Moreover, it includes a selection of <strong><a href="https://www.openair.hosted.york.ac.uk/">Impulse Response (IR) samples</a></strong> with varying Reverberation Time (RT) values, providing a wide range of conditions for evaluation. This data-compilation can be used together with provided Python scripts (available on <strong><a href="http://github.com/TulioChiodi/AIRCADE">GitHub</a></strong>), for generating auralized data ensembles in different sizes: <em>tiny</em>, <em>small</em>, <em>medium</em> and <em>large</em>. Additionally, the provided metadata annotations also allow for further analysis and investigation of the performance of dereverberation algorithms under different conditions. All data is licensed under <strong><a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">Creative Commons Attribution 4.0 International License</a></strong>.</p> <p><strong>About the sizeable versions:</strong></p> <p>The data-compilation is hosted here at <strong><a href="https://zenodo.org/record/7818761#.ZD7ON3bMJPa">Zenodo</a></strong>, with an approximate total file size of 1.3 GB. For simplicity, all samples in our data-compilation were renamed, e.g., <em>guitar_0000</em>, <em>rir_0000</em>, <em>song_0000</em>, <em>speech_0000</em>, and so on. The ensemble versions are available in different sizes, from a <em>tiny</em> version, with limited data, to a <em>large</em> version, with almost 300,000 samples. This allows users to choose the most suitable version for their specific research needs. The following table illustrates the differences between all versions, detailing the number of song, speech, guitar, IR and auralized samples in each one, together with their respective total file size and duration.</p> <table align="center"> <caption>Number of anechoic, IR and resultant auralized data samples, together with their respective total duration and file size for each ensemble version</caption> <tbody> <tr> <td><strong>Version</strong></td> <td><strong>Tiny</strong></td> <td><strong>Small</strong></td> <td><strong>Medium</strong></td> <td><strong>Large</strong></td> </tr> <tr> <td>Song samples</td> <td>100</td> <td>500</td> <td>1,012</td> <td>1,012</td> </tr> <tr> <td>Speech samples</td> <td>100</td> <td>500</td> <td>1,012</td> <td>1,440</td> </tr> <tr> <td>Guitar samples</td> <td>100</td> <td>500</td> <td>1,012</td> <td>2,004</td> </tr> <tr> <td>IR samples</td> <td>5</td> <td>9</td> <td>33</td> <td>65</td> </tr> <tr> <td>Auralized samples</td> <td>1,500</td> <td>13,500</td> <td>100,188</td> <td>289,640</td> </tr> <tr> <td>Total duration</td> <td>3.2 h</td> <td>30.41 h</td> <td>221.77 h</td> <td>658.08 h</td> </tr> <tr> <td>Total file size (required)</td> <td>1.1 GB</td> <td>10.5 GB</td> <td>76.6 GB</td> <td>227.5 GB</td> </tr> </tbody> </table> <p>For more information, please refer to our data paper on <strong><a href="https://arxiv.org/abs/2304.09318">ArXiv</a></strong>.</p> <p><strong>Citation</strong>:</p> <p>If you find <strong>AIRCADE </strong>useful in your research, please cite:</p> <blockquote> <pre>@misc{chiodi2023aircade, title={AIRCADE: an Anechoic and IR Convolution-based Auralization Data-compilation Ensemble}, author={Túlio Chiodi and Arthur dos Santos and Pedro Martins and Bruno Masiero}, year={2023}, eprint={2304.09318}, archivePrefix={arXiv}, primaryClass={eess.AS} }</pre> </blockquote> <p><strong>Acknowledgement</strong>:</p> <p>This work was partially supported by the <strong><a href="https://fapesp.br/">São Paulo Research Foundation (FAPESP)</a></strong>, grants #2017/08120-6 and #2019/22795-1.</p>
Selected BMRB 1D 1H NMR data and physical chemistry values compiled from literature
<p>This dataset contains a collection of a few 1D 1H nuclear magnetic resonance (NMR) spectroscopy experiment data from the Biological Magnetic Resonance Data Bank (BMRB). I collected them for reference on Zenodo because the BMRB in recent years have switched servers and adopted new web APIs, and I want to have this data in a data archive for ease of reproducing the results in my work. Please cite (doi: 10.1093/nar/gkac1050) if you use the BMRB data from this dataset, or consider downloading from their website.</p> <p>This dataset also contains my compiled lists of physical chemistry NMR parameters (chemical shift, J-coupling) from literature and public domain sources for select compounds. One source is the Guided Ideographic Spin System Model Optimization (GISSMO) website, which is based on (DOI: 10.1021/acs.analchem.7b02884) and (DOI: 10.1021/acs.analchem.8b02660). Another source I used is (DOI: 10.1002/nbm.3336). Please cite these sources in addition to this dataset if you use any of the physical chemistry information in this dataset. See the read me file for the format details.</p> <p>I do not guarantee the accuracy of any of the data in this dataset.</p>
Compiled dataset and posterior results for NEMo analyses on Acacia tolerance to aridity and salinity
Open the record for dataset details and reuse information.
Puppet Traces and Compiled Catalogs
<p>This dataset contains the traces and compiled catalogs of the Puppet modules examined in the evaluation of the ICSE'20 paper "Practical Fault Detection in Puppet Programs".</p>
Research Data for article "Clava: C/C++ source-to-source compilation using LARA"
<p>Setup and results for the Section "4. Impact" of the article "Clava: C/C++ source-to-source compilation using LARA"</p> <p>#4.1. Stress Test</p> <p>Instruments several large programs so that they produce a call graph when the program executes.</p> <p>To run the test use the command: clava -c stress_test.clava</p> <p>Some examples (e.g., gcc.c) will only parse sucessfully on a Linux machine.</p> <p>## Files</p> <p>'stats-raw_data.zip' - Raw results taken for the article.</p> <p>'processed_stats.json' - Processed results that are presented in the article.</p> <p><br> #4.3. OpsCounter</p> <p>Instruments the NAS benchmark set so that it counts the number of source code operations executed by the kernels.</p> <p>To run the test use the command: clava -c ops_counter.clava</p> <p>## Files</p> <p>'OpsCounterNAS_S_W_A.json' - Raw results taken for the article.</p> <p>'Results.xlsx' - Processed results that are presented in the article.</p>
Insights into the operation of the solid Earth system from analysis of compiled geochemical data (Video)
<p>This is the first session video recording of the Goldschmidt 2020 Virtual Workshop: Earth Science meets Data Science - Services & Systems, Policies & Procedures, Tools & Techniques for Geochemistry. Moderated by Kerstin Lehnert (Columbia University)</p>
The data generated or compiled in the study of "Diverse polygonal patterned grounds in the northern Eridania basin, Mars: Possible origins and implications"
<p>The data that we generated or compiled, that is used to make a figure (e.g. geologic maps (Figure 2), mapped ridges (included in Figure 2), tables of measurements underlying the histograms (Figures 10 and 18), crater counts (Figure 12)), are placed in a public data repository outside of the JGR paywall. They have been introduced in the following publication, where more details can be found:</p> <p> </p> <p>Y. Dang, F. Zhang, J. Zhao, J. Wang, Y. Xu, T. Huang, and L. Xiao</p> <p>Diverse Polygonal Patterned Grounds in the Northern Eridania Basin, Mars: Possible Origins and Implications.</p> <p>Journal of Geophysical Research: Planets, 2020,125, e2020JE006647. https://<br> doi.org/10.1029/2020JE006647</p> <p> </p> <p>The file format is "Shp" processed in the software ArcGIS 10.6. The detailed description and information can be found in the title of each file below, and also their corresponding figure caption in the paper after it is formally published.</p>
Southern Ocean Cloud and Aerosol data set: a compilation of measurements from the 2018 Southern Ocean Ross Sea Marine Ecosystems and Environment voyage
<p>Due to its remote location and extreme weather conditions, atmospheric in situ measurements are rare in the Southern Ocean. As a result, aerosol-cloud interactions in this region are poorly understood and remain a major source of uncertainty in climate models. This, in turn, contributes substantially to persistent biases in climate model simulations, numerical weather prediction models and reanalyses. It has been shown in previous studies that in situ and ground-based remote sensing measurements across the Southern Ocean are critical for complementing satellite data sets due to the importance of boundary layer and low-level cloud processes. These processes are poorly sampled by satellite-based measurements which are typically obscured by near-continuous overlying cloud cover observed in this region. Here we provide a comprehensive set of ship-based aerosol and meteorological observations collected on the TAN1802 voyage of R/V Tangaroa across the Southern Ocean, from Wellington, New Zealand, to the Ross Sea, Antarctica. The voyage was carried out from 8 February to 21 March, 2018. The compiled data set provides here includes measurements from a range of instruments, such as (i) meteorological conditions at the sea surface and profile measurements; (ii) the size and concentration of particles; (iii) trace gases dissolved in the ocean surface such as dimethyl sulfide and carbonyl sulfide; (iv) and remotely sensed observations of low clouds. We encourage the scientific community to use these measurements for further analysis and model evaluation studies, in particular, for studies of Southern Ocean clouds, aerosol and their interaction.</p>
From Networks of Texts to Networks of Genres? On the Classification of Texts in Compilations with a View towards Manuscript Transmission
<p>At the end of the 14th century, Jakob Twinger von Königshofen, a cleric from Strasbourg, composed a chronicle in the vernacular that spread widely – up to today nearly 130 manuscripts are known that contain the text, wholly or in parts, and that were produced not only in Strasbourg, but as far as Cologne, Augsburg, or Tyrol. About thirty qualify as true copies, while in the big majority of the witnesses, the text is altered in various ways: abbreviated, augmented, updated, corrected, put in a dierent order, and more often than not combined with other texts, either with distinct text boundaries or resulting in new compositions made of several texts.<br> In several manuscripts, a combination of historiographical texts – chronicles, annals, lists, etc. – can be observed, leading to the assumption that Twinger’s work was preferably copied for historiographical compilations and that its structure facilitated some historiographical activity of the recipients. But this view puts a big weight on the genre of Twinger’s work, making it a filter through which all the other texts in a codex are seen. As I have been working on the chronicle transmission, the attribute in common of the known manuscripts is of course the appearance of at least some lines of the Twinger chronicle; but the attempt to look at a single codex as objectively as possible, without giving preference to a particular text, can not only reveal connections between different manuscripts, but also the fluidity and flexibility of medieval texts. The classification of a work as part of a certain genre does not necessarily hold true for its entire transmission, for the manifestation of a distinct text – what is, on the one hand, a nuisance, but on the other a chance to better understand medieval text transmission, manuscript production and the transfer of knowledge.<br> The co-occurrence of certain texts in several manuscripts hints towards intentional copying processes that are reflecting particular interests not only of one individual scribe or commissioner; multiple occurrences of particular compilations can reveal networks that go unseen if the content of a codex is not regarded as a whole. While a text-based analysis often faces difficulties that result from insufficient manuscript descriptions, a broader view that would compare less particular texts, but more areas of interest or fields of knowledge, has to deal with problems regarding the classification of the single texts: Apart from an inevitable subjectivity,<br> questions about the criteria and levels of classification have to be addressed, while the danger of over- or under-representation is always lurking. Discussing these issues and the applicability and usefulness of the comparison of codices from a kind-of-genre-perspective could be fruitful to develop a better understanding for the transmission of manuscripts, of texts and of knowledge.</p>
Compilation of experimental data on water storage capacity in olivine, wadsleyite and ringwoodite (061721 updated)
<p>We compiled high-pressure mineral physics data on water storage capacities in olivine, wadsleyite, and ringwoodite from the literature (published as Tables S1–S3 in Dong et al., 2021). The references are listed in the second tab of each xlsx file.</p> <p>This version was published on Oct. 21, 2020 (version 1.0) and includes "data-olivine-table-s1.xlsx", "data-wadsleyite-table-s2.xlsx", and "data-ringwoodite-table-s3.xlsx"</p> <p>Please cite "Dong, J., Fischer, R. A., Stixrude, L. P., & Lithgow‐Bertelloni, C. R. (2021). Constraining the Volume of Earth's Early Oceans With a Temperature‐Dependent Mantle Water Storage Capacity Model. <em>AGU Advances</em>, <em>2</em>(1), e2020AV000323." for the dataset version 1.0.</p> <p><strong>An update to the data compilation of Dong et al. (2021) (version 1.0) is now published as version 2.0 on Jun. 17, 2021, and includes additional literature data on the water storage capacity in olivine, wadsleyite, and ringwoodite ("additional-data-for-mars-061721.xlsx")</strong></p>
Compiler Equivalence
<p>Identifying equivalent mutants remains the largest impediment to the widespread uptake of mutation testing. Despite being researched for more than three decades, the problem remains. We propose Trivial Compiler Equivalence (TCE) a technique that exploits the use of readily available compiler technology to address this long-standing challenge. TCE is directly applicable to real-world programs and can imbue existing tools with the ability to detect equivalent mutants and a special form of useless mutants called duplicated mutants. We present a thorough empirical study using 6 large open source programs, several orders of magnitude larger than those used in previous work, and 18 benchmark programs with hand-analysis equivalent mutants. Our results reveal that, on large real-world programs, TCE can discard more than 7% and 21% of all the mutants as being equivalent and duplicated mutants respectively. A human-based equivalence verification reveals that TCE has the ability to detect approximately 30% of all the existing equivalent mutants.</p>
Figure 2. - Number of species recorded in Jean Gutierrez collection dataset (solid bar) and in the literature (dashed bar) compiled in Spider Mites Web (http://www1.montpellier.inra.fr/CBGP/spmweb/) for the areas of particular interest. Colour scheme same as in Figure 1.
Figure 2. - Number of species recorded in Jean Gutierrez collection dataset (solid bar) and in the literature (dashed bar) compiled in Spider Mites Web (http://www1.montpellier.inra.fr/CBGP/spmweb/) for the areas of particular interest. Colour scheme same as in Figure 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.