Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,505
datasets available to search
ShareScore release 0.7.1
Dataset results
7,505 results for “Generation”
Fig. 7 in Unexpected species diversity in electric eels with a description of the strongest living bioelectricity generator
Fig. 7 Lateral view of Electrophorus voltai sp. nov. Holotype, Museu Paraense Emílio Goeldi MPEG 15529, 1290 mm TL. Ipitinga River, Brazil
Fig. 6 in Unexpected species diversity in electric eels with a description of the strongest living bioelectricity generator
Fig. 6 Lateral view of Electrophorus varii sp. nov. Holotype, Museu Paraense Emílio Goeldi MPEG 25422, 1000 mm TL. Goiapi River, Brazil
Test Suites from Test-Generation Tools (Test-Comp 2021)
<p>Test Suites</p> <p>This file describes the contents of an archive of the 3rd Competition on Software Testing (Test-Comp 2021).<br> <a href="https://test-comp.sosy-lab.org/2021/">https://test-comp.sosy-lab.org/2021/</a></p> <p>The competition was run by Dirk Beyer, LMU Munich, Germany.<br> More information is available in the following article:<br> Dirk Beyer. <em>Status Report on Software Testing: Test-Comp 2021.</em> In Proceedings of the 24th International Conference on Fundamental Approaches to Software Engineering (FASE 2021, Luxembourg, March 27 - April 1), 2021. Springer.</p> <p>Copyright (C) Dirk Beyer<br> <a href="https://www.sosy-lab.org/people/beyer/">https://www.sosy-lab.org/people/beyer/</a></p> <p>SPDX-License-Identifier: CC-BY-4.0<br> <a href="https://spdx.org/licenses/CC-BY-4.0.html">https://spdx.org/licenses/CC-BY-4.0.html</a></p> <p>Contents</p> <ul> <li><code>LICENSE.txt</code>: specifies the license</li> <li><code>README.txt</code>: this file</li> <li><code>witnessFileByHash/</code>: This directory contains test suites (witnesses for coverage). Each witness in this directory is stored in a file whose name is the SHA2 256-bit hash of its contents followed by the filename extension .zip. The format of each test suite is described on the format web page: <a href="https://gitlab.com/sosy-lab/software/test-format">https://gitlab.com/sosy-lab/software/test-format</a> A test suite contains also metadata in order to relate it to the test problem for which it was produced.</li> <li><code>witnessInfoByHash/</code>: This directory contains for each test suite (witness) in directory witnessFileByHash/ a record in JSON format (also using the SHA2 256-bit hash of the witness as filename, with .json as filename extension) that contains the meta data.</li> <li><code>witnessListByProgramHashJSON/</code>: For convenient access to all test suites for a certain program, this directory represents a function that maps each program (via its SHA2256-bit hash) to a set of test suites (JSON records for test suites as described above) that the test tools have produced for that program. For each program for which test suites exist, the directory contains a JSON file (using the SHA2 256-bit hash of the program as filename, with .json as filename extension) that contains all JSON records for test suites for that program.</li> </ul> <p>A similar data structure was used by SV-COMP and is described in the following article:<br> Dirk Beyer. <em>A Data Set of Program Invariants and Error Paths.</em> In Proceedings of the 2019 IEEE/ACM 16th International Conference on Mining Software Repositories (MSR 2019, Montreal, Canada, May 26-27), pages 111-115, 2019. IEEE.<br> <a href="https://doi.org/10.1109/MSR.2019.00026">https://doi.org/10.1109/MSR.2019.00026</a></p> <p>Other Archives</p> <p>Overview over archives from Test-Comp 2021 that are available at Zenodo:</p> <ul> <li><a href="https://doi.org/10.5281/zenodo.4459466">https://doi.org/10.5281/zenodo.4459466</a> Witness store (containing the generated test suites)</li> <li><a href="https://doi.org/10.5281/zenodo.4459470">https://doi.org/10.5281/zenodo.4459470</a> Results (XML result files, log files, file mappings, HTML tables)</li> <li><a href="https://doi.org/10.5281/zenodo.4459132">https://doi.org/10.5281/zenodo.4459132</a> Test tasks, version testcomp21</li> <li><a href="https://doi.org/10.5281/zenodo.4317433">https://doi.org/10.5281/zenodo.4317433</a> BenchExec, version 3.6</li> </ul> <p>All benchmarks were executed for Test-Comp 2021 <a href="https://test-comp.sosy-lab.org/2021/">https://test-comp.sosy-lab.org/2021/</a><br> by Dirk Beyer, LMU Munich, based on the following components:</p> <ul> <li><a href="https://gitlab.com/sosy-lab/test-comp/archives-2021">https://gitlab.com/sosy-lab/test-comp/archives-2021</a> testcomp21-0-gdacd4bf</li> <li><a href="https://gitlab.com/sosy-lab/software/sv-benchmarks">https://gitlab.com/sosy-lab/software/sv-benchmarks</a> testcomp21-0-gefea738258</li> <li><a href="https://gitlab.com/sosy-lab/software/benchexec">https://gitlab.com/sosy-lab/software/benchexec</a> 3.6-0-gb278ebbb</li> <li><a href="https://gitlab.com/sosy-lab/benchmarking/competition-scripts">https://gitlab.com/sosy-lab/benchmarking/competition-scripts</a> testcomp21-0-g8339740</li> <li><a href="https://gitlab.com/sosy-lab/test-comp/bench-defs">https://gitlab.com/sosy-lab/test-comp/bench-defs</a> testcomp21-0-g9d532c9</li> </ul> <p>Contact</p> <p>Feel free to contact me in case of questions: <a href="https://www.sosy-lab.org/people/beyer/">https://www.sosy-lab.org/people/beyer/</a></p>
A set of generated Instagram Data Download Packages (DDPs) to investigate their structure and content
<p><strong>Instagram data-download example dataset</strong></p> <p>In this repository you can find a data-set consisting of 11 personal Instagram archives, or Data-Download Packages (DDPs).</p> <p> </p> <p><strong>How the data was generated</strong></p> <p>These Instagram accounts were all new and generated by a group of researchers who were interested to figure out in detail<br> the structure and variety in structure of these Instagram DDPs. The participants user the Instagram account extensively for approximately a week. The participants also intensively communicated with each other so that the data can be used as an example of a network. </p> <p>The data was primarily generated to evaluate the performance of de-identification software. Therefore, the text in the DDPs particularly contain many randomly chosen (Dutch) first names, phone numbers, e-mail addresses and URLS. In addition, the images in the DDPs contain many faces and text as well. The DDPs contain faces and text (usernames) of third parties. However, only content of so-called `professional accounts' are shared, such as accounts of famous individuals or institutions who self-consciously and actively seek publicity, and these sources are easily publicly available. Furthermore, the DDPs do not contain sensitive personal data of these individuals. </p> <p><br> <strong>Obtaining your Instagram DDP</strong></p> <p>After using the Instagram accounts intensively for approximately a week, the participants requested their personal Instagram DDPs by using the following steps. You can follow these steps yourself if you are interested in your personal Instagram DDP. </p> <p>1. Go to www.instagram.com and log in<br> 2. Click on your profile picture, go to *Settings* and *Privacy and Security*<br> 3. Scroll to *Data download* and click *Request download*<br> 4. Enter your email adress and click *Next*<br> 5. Enter your password and click *Request download*</p> <p>Instagram then delivered the data in a compressed zip folder with the format **username_YYYYMMDD.zip** (i.e., Instagram handle and date of download) to the participant, and the participants shared these DDPs with us.</p> <p> </p> <p><strong>Data cleaning</strong></p> <p>To comply with the Instagram user agreement, participants shared their full name, phone number and e-mail address. In addition, Instagram logged the i.p. addresses the participant used during their active period on Instagram. After colleting the DDPs, we manually replaced such information with random replacements such that the DDps shared here do not contain any personal data of the participants.</p> <p> </p> <p><strong>How this data-set can be used</strong></p> <p>This data-set was generated with the intention to evaluate the performance of the de-identification software. We invite other researchers to use this data-set for example to investigate what type of data can be found in Instagram DDPs or to investigate the structure of Instagram DDPs. The packages can also be used for example data-analyses, although no substantive research questions can be answered using this data as the data does not reflect how research subjects behave `in the wild'. </p> <p><br> <strong>Authors</strong></p> <p>The data collection is executed by Laura Boeschoten, Ruben van den Goorbergh and Daniel Oberski of Utrecht University. For questions, please contact l.boeschoten@uu.nl. </p> <p> </p> <p><strong>Acknowledgments</strong></p> <p>The researchers would like to thank everyone who participated in this data-generation project.</p>
Unitymol demo on generating 360 degree videos
<p>This video provides more detailed supportive information about using Unitymol.</p> <p>It requires to run UnityMol from the Unity editor. It will not work with the UnityMol executable.</p> <p>1) from within the Unity editor with an opened UnityMol project, open the Recorder window<br> 2) add a new Movie recorder, setting parameters as shown in the tutorial video<br> 3) launch the UnityMol project (play mode) and start recording<br> 4) do what you want to be captured within UnityMol (can be live action, can be executing a pre-recorded script ..)<br> 5) stop recording when finished</p>
Evaluation results for When a Computer Cracks a Joke: Automated Generation of Humorous Headlines
<p>Evaluation results for the paper:</p> <p>Alnajjar, K., & Hämäläinen, M. (2021) When a Computer Cracks a Joke: Automated Generation of Humorous Headlines. In<em> The Proceedings of the Twelfth International Conference on Computational Creativity, ICCC’21</em>.</p> <p>The table has the aggregated evaluation results for each evaluation question. The left and right columns are the title before and after the replacement word. The replacement column show the humorous word and original column the word that existed in the headline before the replacement. The system column indicates whether the humorous headline was produced by our system or by a human.</p>
North Atlantic Oscillation (NAO) climate index hidden in ocean generated secondary microseisms
<p>Datatsets associated with "North Atlantic Oscillation (NAO) climate index hidden in ocean generated secondary microseisms". The data include the daily seismic cross-correlograms for station pairs located on land and at the seafloor offshore Ireland, 3D models used for the numerical simulations and the associated synthetic seismic data.</p>
Dataset: On the Properties of Next Generation Wireless Backhaul
<p>This dataset contains all the data used in the Paper: "On the Properties of Next Generation Wireless Backhaul" sent to Transaction of Network Science and Engineering.</p> <p>It is divided into three main archives:</p> <ul> <li>The first archive, called <strong>geodata.zip</strong>, contains the Data Surface Model (DSM) and the topographical maps used to generate the intervisibility graphs. These maps have been aggregated from different sources and they are all released under a CC-BY-SA 4.0 license. It is divided into two folders: 'dsm' and 'topo_maps': <ul> <li> 'dsm' contains the Data Surface Model (DSM) of the nine areas used in our research. Each area is saved in a separate .tif file and can be used directly from the tool. The projection system is EPSG:3003.</li> <li>'topo_mapsì contains the PostGIS dump of the tables containing the topographical maps of the cities. In order to use them, you have to import those in PostGIS and connect our tool to the PostGIS DB.<br> The CTR files (technical region maps) are divided by region, while the openstreetmap file (osm.tar) contains the whole Italian peninsula. All the geographical data are projected in the reference system EPSG:3003. The CTR of Campania is unavailable due to licensing incompatibilities.</li> </ul> </li> <li>The second archive, called <strong>visibility_graphs.zip</strong>, contains the visibility graph we have generated using our tool. It is released as a CC-BY-SA 4.0 license. It is divided into folders, one for each area of the analysis. Each folder contains two files: <ul> <li>best_p.csv : contains the ids of the nodes associated with their coordinates on a cartesian projection. The projection is EPSG:3003</li> <li>intervisibility.adj : contains the intervisibility graph represented as an adjacency matrix. It can be read by common libraries such as networkx or igraph. The intervisibility has been calculated 2m above the building.</li> </ul> </li> <li>The third folder, called <strong>network_topologies.zip</strong>, contains the network topologies we have computed using a State-of-the-Art algorithm from Polese et Al. The folder is divided into two folders: '30gNB' and '60gNB', corresponding to the density of base stations for square km. It s further divided in 9 folders (one for each area), then divided on the basis of the model we use for the visibility graph. You can find more information about these models in section V of the paper. In this last folder we have two sets of 50 files, which are the intervisibility graph of that set of Base Stations and the topology produced. This dataset has been released under a CC-BY-SA 4.0 license.</li> </ul> <p>All the results can be replicated using our code which will be released with an open-source license as soon as the article is published.</p>
Photovoltaic generation and temperature for the year 2019
<p>This dataset has photovoltaic generation data and temperature data regarding a research building in ISEP/P.Porto (Instituto Superior de Engenharia do Porto / Politécnico do Porto). The data was measured using 5-minutes periods during the entire year of 2019. The temperature sensor was located near the photovoltaic panels (without having direct sunlight). The photovoltaic installation has a theoretical peak generation of 7.5 kW.<br>The dataset presents some errors in the data, representing failures in the acquisition system.</p> <p> </p> <p>We would be grateful if you could acknowledge the use of this dataset in your publications. Please use the Zenodo publication to cite this work.</p>
Supporting dataset for manuscript: "Higher rate of tuberculosis in second generation migrants compared to native residents in a metropolitan setting in Western Europe" (PLoS ONE)
<p>This is the supporting datafile for the manuscript entitled "Higher rate of tuberculosis in second generation migrants compared to native residents in a metropolitan setting in Western Europe" (Marx et al., PLoS ONE). The dataset includes anonymized, routinely collected notification data (variables labeled as "nd") for 314 individuals and anonymized survey data (i.e. data obtained through interviews; variables labeled as "sd") for a subset of 154 individuals. The data are published open-access, in accordance with the PLoS ONE data policy (2014).</p>
Raw NGS data for the study 'Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing'
<p>This file contains the original next-generation sequencing data (raw sequences in fastq format) which were analyzed in the study titled: "Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing".</p> <p> </p> <p> </p>
Graphing and tabulating next-generation sequencing and genotyping data
<p>Making figures and tables for publication. Each zip archive contains input data, shell script to initiate and log R script, one R script for generating several graphs and tables, and the output graphs and tables themselves.</p> <p>Data was generated by whole-genome resequencing of 22 individual D.melanogaster from Sussex-LHM population and 2 from the Sussex RG line, followed by read-mapping, then genotyping with Haplotype Caller and Genomestrip.</p> <p>Locations for raw data, code, logs, extended QC data:</p> <p>Sequence reads NCBI SRA268956</p> <p>NCBI dbSNP https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461</p> <p>NCBI dbVar accession number pre-release nstd134</p> <p> </p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>
Read-mapping for next-generation sequencing data (Drosophila melanogaster)
<p>Code, logs and quality-control data for whole-genome resequencing of Sussex-LH<sub>M</sub> and RG <em>Drosophila melanogaster</em>.</p> <p>Mapping code is in the archive lhm_mapping_scripts.zip</p> <p>Mapping logs are in the in the archive lhm_mapping_logs.zip</p> <p>Other zip archives contain the quality control data.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Read-mapping for next-generation sequencing data (Wolbachia)
<p>Code, log files and QC data. NCBI SRA accession number SRP091004. Note that sequencing Wolbachia was not a central aim of the project, and was undertaken in order to maximise the amount of information that could be extracted from the raw genome sequence data targetted at the fruit-fly host.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Genotype reproducibility testing in next-generation sequencing data
<p>Code, log and results summary for testing the reproducibility of genotypes with three pairs of hemiclones in the Sussex LH<sub>M </sub><em>D.melanogaster </em>population sample. Discovery and genotyping of genomic sequence variants was done using GATK HaplotypeCaller, and Genomestrip. Numerical comparison of genotype calls within each pairs of hemiclone individuals was performed using GATK GenotypeConcordance.</p> <p> </p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Two datasets with user generated audio recordings
<p>We provide two open access datasets of <strong>user generated audio recordings </strong>captured with mobile devices such as smartphones and portable cameras . The provided audio files originate from two different public events, a <strong>musical concert</strong> and a <strong>football match</strong>. Also, for each event, we provide two different types of collections; the original <strong>unorganized</strong> collection of uncompressed audio files and an additional <strong>organized</strong> collection, where the different recordings corresponding to similar parts of the event are grouped into specific folders and time-aligned so they can be played back in unison. The interested researcher is invited to read the accompanying paper "<em>Two open access datasets of user generated audio recordings</em>" for finding out more details about these datasets and the way that they can be useful in the context of research related to the organization and reproduction of user generated content.</p>
Test Data Generation from Business Rules
<p><strong>Overview of Data</strong></p> <p>The site includes data only for the two subjects: Ceu-pacific and JBilling. For both the subjects, the “<em>.model” shows the model created from the business rules obtained from respective websites, and “</em>_HighLevelTests.csv” shows the tests generated. Among csv files, we show tests generated by both BUSTER and Exhaust as well.</p> <p><strong>Paper Abstract</strong></p> <p>Test cases that drive an application under test via its graphical user interface (GUI) consist of sequences of steps that perform actions on, or verify the state of, the application user interface. Such tests can be hard to maintain, especially if they are not properly modularized—that is, common steps occur in many test cases, which can make test maintenance cumbersome and expensive. Performing modularization manually can take up considerable human effort. To address this, we present an automated approach for modularizing GUI test cases. Our approach consists of multiple phases. In the first phase, it analyzes individual test cases to partition test steps into candidate subroutines, based on how user-interface elements are accessed in the steps. This phase can analyze the test cases only or also leverage execution traces of the tests, which involves a cost-accuracy tradeoff. In the second phase, the technique compares candidate subroutines across test cases, and refines them to compute the final set of subroutines. In the last phase, it creates callable subroutines, with parameterized data and control flow, and refactors the original tests to call the subroutines with context-specific data and control parameters. Our empirical results, collected using open-source applications, illustrate the effectiveness of the approach.</p>
Data for article: Time-Resolved Spectroscopic Investigation of Charge Trapping in Carbon Nitrides Photocatalysts for Hydrogen Generation
<p>This is the data presented in the article titled 'Time-Resolved Spectroscopic Investigation of Charge Trapping in Carbon Nitrides Photocatalysts for Hydrogen Generation', published in the Journal of the American Chemical Society. DOI:10.1021/jacs.7b01547</p> <p>http://pubs.acs.org/doi/abs/10.1021/jacs.7b01547</p> <p> </p>
An Open Access User Generated Video Dataset from 2016 Edinburgh Festival
<p>A user generated video dataset captured during the 2016 Edinburgh festival. The provided dataset was collected using a smart phone and is available with no post-processing. The videos mainly cover the Edinburgh streets and the festival atmosphere, and do not cover any performances. The dataset can be used for evaluation of various research tools, such as video quality assessment and enhancement.</p>
Evaluation of two 4th generation point-of-care assays for the detection of Human Immunodeficiency Virus infection.
<p>Background. Fourth generation assays detect simultaneously antibodies for HIV and the p24 antigen, identifying HIV infection earlier than previous generation tests. Previous studies have shown that the Alere Determine HIV-1/2 Combo has lower than anticipated performance in detecting antibodies for HIV and the p24 antigen. Furthermore, there are currently very few studies evaluating the performance of Standard Diagnostics BIOLINE HIV Ag/Ab Combo.</p> <p>Objective: To evaluate the performance of the Alere Determine HIV-1/2 Combo and the Standard Diagnostics BIOLINE HIV Ag/Ab Combo in a panel of frozen serum samples.</p> <p>Study Design: The testing panel included 133 previously frozen serum specimens from the UCLA Clinical Microbiology & Immunoserology laboratory. Reference testing included testing for HIV antibodies by a 3<sup>rd</sup> generation enzyme immunoassay followed by HIV RNA detection. Antibody negative and RNA positive sera were also tested by a laboratory 4<sup>th</sup> generation HIV Ab/Ag enzyme immunoassay.</p> <p>Results: Reference testing yielded 97 positives for HIV infection and 36 negative samples. Sensitivity of the Alere test was 95% (88-98%), while the SD Bioline sensitivity was 91% (83-96%). Both assays showed 100% (90-100%) specificity. No indeterminate or invalid results were recorded. Among 13 samples with acute infection (HIV RNA positive, HIV antibody negative), 12 were found positive by the first assay and 8 by the second. The antigen component of the Alere assay detected 10 acute samples, while the SD Bioline assay detected only one.</p> <p>Conclusions: Both rapid assays showed very good overall performance in detecting HIV infection in frozen serum samples, but further improvements are required to improve the performance in acute infection.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.