Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
915
datasets available to search
ShareScore release 0.9.0
Dataset results
915 results for “Graph”
Is this bug severe? A text-cum-graph based model for bug severity prediction
<p>A snapshot of the dataset has been updated. For the time being, we are publishing a snapshot of the dataset where the bugs were reported after 2017.</p> <p>Paper link: <a href="https://arxiv.org/abs/2207.00623">https://arxiv.org/abs/2207.00623</a> (ECML-PKDD 2022)</p> <p>Cite our paper:</p> <p>@InProceedings{10.1007/978-3-031-26422-1_15,<br> author="Hazra, Rima<br> and Dwivedi, Arpit<br> and Mukherjee, Animesh",<br> editor="Amini, Massih-Reza<br> and Canu, St{\'e}phane<br> and Fischer, Asja<br> and Guns, Tias<br> and Kralj Novak, Petra<br> and Tsoumakas, Grigorios",<br> title="Is This Bug Severe? A Text-Cum-Graph Based Model for Bug Severity Prediction",<br> booktitle="Machine Learning and Knowledge Discovery in Databases",<br> year="2023",<br> publisher="Springer Nature Switzerland",<br> address="Cham",<br> pages="236--252",<br> isbn="978-3-031-26422-1"<br> }</p> <p><strong>*** Please see the new version. (10.5281/zenodo.5554978)</strong></p> <p>There is a total of six files.</p> <ul> <li><strong>bug_descriptions.csv:</strong> This file contains the bug id and its description.</li> <li><strong>bug_comments.csv:</strong> This file contains three columns. The columns are the bug ids, comments and timestamp of the comment.</li> <li><strong>bug_REPORTED_ON_details.csv:</strong> This file contains the bug id and the package name on which the bug is reported</li> <li><strong>affect_dataset.csv: </strong>This file contains the bug id and the affected packages along with the affect timestamp.</li> <li><strong>bug_heat_2019.csv:</strong> This file contains the bug ids and its bug heats crawled in November 2019.</li> <li><strong>bug_heat_2020.csv:</strong> This file contains the bug ids and its bug heats crawled in November 2020.</li> </ul>
Universal Knowledge Graph Embeddings
<p>The dataset provides embeddings for entities and relations in DBpedia (English) and Wikidata. The two knowledge graphs are first merged using a novel approach that we developed by leveraging the sameAs links between them. Then, we used the state-of-the-art embedding model ConEx to compute embeddings of the merge. Our embeddings are called universal knowledge graph embeddings.</p>
Toloker Graph: Interaction of Crowd Annotators
<p>The graph contains 11,758 nodes and 519,000 edges representing interactions between crowd annotators on a project labeled on the <a href="https://toloka.ai/">Toloka</a> crowdsourcing platform (see the <a href="https://toloka.ai/en/docs/guide/concepts/overview">Toloka overview</a> for the details on the used terminology).</p> <p>Each node represents an individual annotator; nodes are provided with four numerical and three categorical features. An edge is drawn between a pair of annotators if they annotated the same task. Also, each node is provided with a label showing whether the annotator was banned on this project, or not.</p> <p><strong>Nodes</strong> are stored in the <a href="https://github.com/Toloka/TolokerGraph/blob/main/nodes.tsv">nodes.tsv</a> file in the TSV format of the following structure:</p> <ul> <li><code>id</code>: unique identifier of the annotator</li> <li><code>approved_rate</code>: percentage of the approved labels of this annotator</li> <li><code>skipped_rate</code>: percentage of the skipped tasks of this annotator</li> <li><code>expired_rate</code>: percentage of the expired tasks of this annotator</li> <li><code>rejected_rate</code>: percentage of the rejected labels of this annotator</li> <li><code>education</code>: level of education as self-reported by this annotator (<code>none</code>, <code>basic</code>, <code>middle</code>, <code>high</code>)</li> <li><code>english_profile</code>: knowledge of English as self-reported by this annotator (<code>0</code> for no, <code>1</code> for yes)</li> <li><code>english_tested</code>: whether the annotator passed the Toloka language test for English (<code>0</code> for no, <code>1</code> for yes)</li> <li><code>banned</code>: whether the annotator was banned on this project (<code>0</code> for no, <code>1</code> for yes)</li> </ul> <p>The <code>*_rate</code> attributes should sum up to 1.</p> <p><strong>Edges</strong> are stored in the <a href="https://github.com/Toloka/TolokerGraph/blob/main/edges.tsv">edges.tsv</a> file in the TSV format of the following structure:</p> <ul> <li><code>source</code>: source identifier of the annotator</li> <li><code>target</code>: target identifier of the annotator</li> </ul> <p>As the graph is undirected, <code>source</code> and <code>target</code> can be interchanged for the given pair of nodes.</p>
cultural-ai/wordsmatter: Words Matter: a knowledge graph of contentious terms
<p>The choice of words describing cultural heritage can cause debates. It is especially sensitive when artefacts relate to different cultures and peoples who have been historically marginalised. Words chosen by archivists or curators may transmit stereotypes. The cultural heritage community has produced knowledge on potentially stereotyping and offensive terminology in heritage collections. At the same time, their knowledge is difficult to incorporate into existing online collections unless this knowledge is structured and machine-readable.</p> <p>The Words Matter Knowledge Graph represents domain expert knowledge on discussions about contentious terminology in the cultural sector. In the knowledge graph, 75 English and 83 Dutch contentious terms are linked to explanations of their usage and suggested alternatives from domain experts. There are also related matches between contentious terms and sources from external datasets: Wikidata, Princeton WordNet, Open Dutch WordNet, and Getty Art & Architecture Thesaurus.</p> <p>This Zenodo publication includes the CULCO scheme used to model contentious terms in the knowledge graph. The scheme documentation is <a href="https://cultural-ai.github.io/wordsmatter/" target="_blank" rel="noopener">available on a separate page</a>.</p> <p>This knowledge graph is <a href="https://amsterdam.wereldmuseum.nl/en/about-wereldmuseum-amsterdam/research/words-matter-publication" target="_blank" rel="noopener">based</a> on the publication “Words Matter: An Unfinished Guide to Word Choices in the Cultural Sector” by the National Museum of World Cultures (NMVW). </p> <p><a href="https://doi.org/10.1007/978-3-031-33455-9_30" target="_blank" rel="noopener">Read more</a> about this work in the paper "A Knowledge Graph of Contentious Terminology for Inclusive Representation of Cultural Heritage" (2023) by Andrei Nesterov, Laura Hollink, Marieke van Erp & Jacco van Ossenbruggen.</p> <p>In this version:</p> <ul> <li>the CULCO scheme documentation is updated</li> <li>versioning is fixed</li> <li>typos are corrected</li> </ul>
Graph 4: Publication Types for Heiner Müller's Writings 1978-1988
<p>Graph 4 shows the publication types that Heiner Müller used to publish his writing in the years from 1978-1988.</p>
Graph 2: Publication Type and Venue of Heiner Müller's Writings
<p>This graph shows the different publication types and publication venues that Heiner Müller used to publish his writings (essays, articles, speeches, etc.) from 1968 until 1998.</p>
Dataset Graph 1: Heiner Müller's Publications 1960–1998.
<p>Dataset for Graph 1. Data collected based on the Heiner Müller Werkausgabe.</p> <p>Graph 1 depicts Müller’s publication record regarding his public-oriented writings from the early stages of his career in the 1960s until two years after his death in 1998.</p>
Graph 6: Stagings of Heiner Müller's Plays in Germany (East and West).
<p>Graph 6 shows the number of stagings of Heiner Müller’s plays in East and West Germany from 1957-1990.</p>
Graph 3: Publication Types for Heiner Müller's Writings 1967–1977
<p>Graph 3 shows the publication types that Heiner Müller used to publish his writing in the years from 1967–1977.</p>
Graph 5: Publication Types for Heiner Müller's Writings 1989-1998.
<p>Graph 5 shows the publication types that Heiner Müller used to publish his writing in the years from 1989-1998.</p>
Dataset Graph 2: Publication Type and Venue of Heiner Müller's Writings
<p>Dataset for Graph 2. Data collected based on the Heiner Müller Werkausgabe.</p> <p>Graph 2 shows the different publication types and publication venues that Heiner Müller used to publish his writings (essays, articles, speeches, etc.) from 1968 until 1998.</p>
Graph 1: Heiner Müller's Publications 1960–1998.
<p>Graph 1 depicts Müller’s publication record regarding his public-oriented writings from the early stages of his career in the 1960s until two years after his death in 1998.</p>
Dataset Graph 8: Heiner Müller's Interview Partners
<p>Dataset for Graph 8. Data collected based on the Heiner Müller Werkausgabe.</p> <p>Graph 8 shows Heiner Müller's interview partners and the years in which his interviews were published.</p>
Graph 8: Heiner Müller's Interview Partners
<p>Graph 8 shows Heiner Müller's interview partners and the years in which his interviews were published.</p>
Graph 9: TV and Radio Broadcasts of Heiner Müller's Interviews.
<p>Graph 9 shows when and where Heiner Müller's interviews were initially broadcasted and with whom he recorded the respective interview.</p>
Dataset Graph 9: TV and Radio Broadcasts of Heiner Müller's Interviews.
<p>Dataset for Graph 9. Data collected based on the Heiner Müller Werkausgabe.</p> <p>Graph 9 shows when and where Heiner Müller's interviews were initially broadcasted and with whom he recorded the respective interview.</p>
Apis mellifera graph genome
<p><strong>AmelGraph 1.1.0</strong></p> <p>We aligned 5 different publicly available assemblies with the cactus pangenome workflow (v2.0.5):</p> <table> <tbody> <tr> <td> <p><strong>Use </strong></p> </td> <td> <p><strong>Species </strong></p> </td> <td> <p><strong>Genome ID </strong></p> </td> <td> <p><strong>Accession </strong></p> </td> <td> <p><strong>Graph ID </strong></p> </td> <td> <p><strong>Size (Mb) </strong></p> </td> </tr> <tr> <td> <p>Reference </p> </td> <td> <p><em>A. mellifera</em> (DH4) </p> </td> <td> <p>Amel_HAv3.1 </p> </td> <td> <p>GCF_003254395.2 </p> </td> <td> <p>DH4 </p> </td> <td> <p>225.2 </p> </td> </tr> <tr> <td> <p>Derivate </p> </td> <td> <p><em>A. m. mellifera </em></p> </td> <td> <p>INRA_AMelMel_1.0 </p> </td> <td> <p>GCA_003314205.1 </p> </td> <td> <p>mellifera </p> </td> <td> <p>227.0 </p> </td> </tr> <tr> <td> <p>Derivate </p> </td> <td> <p><em>A. m. carnica </em></p> </td> <td> <p>ASM1384124v2 </p> </td> <td> <p>GCA_013841245.2 </p> </td> <td> <p>carnica </p> </td> <td> <p>226.0 </p> </td> </tr> <tr> <td> <p>Derivate </p> </td> <td> <p><em>A. m. caucasica </em></p> </td> <td> <p>ASM1384120v1 </p> </td> <td> <p>GCA_013841205.1 </p> </td> <td> <p>caucasica </p> </td> <td> <p>224.8 </p> </td> </tr> <tr> <td> <p>Derivate </p> </td> <td> <p><em>A. m. ligustica </em></p> </td> <td> <p>ASM1932182v1 </p> </td> <td> <p>GCA_019321825.1 </p> </td> <td> <p>ligustica </p> </td> <td> <p>231.1 </p> </td> </tr> </tbody> </table> <p> </p> <p>We used the cactus<sup>1</sup> <a href="https://github.com/ComparativeGenomicsToolkit/cactus/blob/master/doc/pangenome.md">pangenome workflow</a><a href="https://github.com/ComparativeGenomicsToolkit/cactus/blob/master/doc/pangenome.md"> </a>to generate a 5-ways pangenome alignment, masking in blocks of 10Kb, and considering all the sequences in the different data. Due to incompatibilities with the pangenome workflow generation of the indexes, we regenerated the giraffe indexes from the output GFA/VCF file. We also downloaded annotations available for three genomes (DH4, <em>A. m. carnica</em> and <em>A. m. caucasica</em>), modifying the naming of the contigs to include the subspecies name (i.e. 'LG1' for <em>A. m. caucasica</em> modified to 'caucasica.LG1'). We then ran the vg autoindex function:</p> <pre><code>vg autoindex -w giraffe -o pangenome -t 8 -T ./TMP -x carnica.gff -x caucasica.gff -R XG -x DH4.gff</code></pre> <p>We compared this 'full' cactus graph to one built using only sequence on the linkage groups, and to one generated using the <a href="https://github.com/pangenome/pggb">PGGB workflow</a>. We evaluated the different graphs using <a href="https://github.com/pangenome/pgge">PGGE</a>. This full cactus graph returned the highst aligned identity, sequence matches, and unique alignments, while having the lowest number of multiple mapping and missing alignments.</p> <p> </p> <table> <tbody> <tr> <td> <p><strong>Parameter </strong></p> </td> <td> <p><strong>CACTUS (LG) </strong></p> </td> <td> <p><strong>CACTUS (FULL) </strong></p> </td> <td> <p><strong>PGGB (norm) </strong></p> </td> </tr> <tr> <td> <p>HAv3.1 size </p> </td> <td> <p>221,626,419 </p> </td> <td> <p>225,250,884 </p> </td> <td> <p>225,250,884 </p> </td> </tr> <tr> <td> <p>Graph length (bp) </p> </td> <td> <p>243,077,200 </p> </td> <td> <p>246,675,144 </p> </td> <td> <p>315,672,860 </p> </td> </tr> <tr> <td> <p>Extra sequence </p> </td> <td> <p>21,450,781 </p> </td> <td> <p>21,424,260 </p> </td> <td> <p>90,421,976 </p> </td> </tr> <tr> <td> <p># nodes </p> </td> <td> <p>17,120,383 </p> </td> <td> <p>11,664,080 </p> </td> <td> <p>17,952,597 </p> </td> </tr> <tr> <td> <p># edges </p> </td> <td> <p>21,385,634 </p> </td> <td> <p>15,990,593 </p> </td> <td> <p>21,702,958 </p> </td> </tr> </tbody> </table> <p> </p> <p><strong>References</strong></p> <p>1. Armstrong, J. <em>et al</em>. Progressive Cactus is a multiple-genome aligner for the thousand-genome era. <em>Nature</em> <strong>587</strong>, (2020). </p> <p> </p>
Graph construction method impacts variation representation and analyses in a bovine super-pangenome
<p>Pangenomes for minigraph, pggb, and cactus containing assemblies from</p> <ul> <li>Hereford (cattle reference genome)</li> <li>Highland</li> <li>Brown Swiss</li> <li>Angus</li> <li>Simmental</li> <li>Original Braunvieh</li> <li>Piedmontese</li> <li>Nellore</li> <li>Brahman</li> <li>Yak</li> <li>Bison</li> <li>Gaur</li> </ul> <p> </p> <p>Also contains the genomic region classifications for the ARS-UCD1.2 reference genome for</p> <ul> <li>Satellites</li> <li>Tandem repeats</li> <li>Low mappability</li> <li>Repetitive regions</li> <li>Normal (everything else)</li> </ul>
LauNuts: A Knowledge Graph to identify and compare geographic regions in the European Union
<p><strong>LauNuts</strong> is a RDF Knowledge Graph consisting of:</p> <ul> <li>Local Administrative Units (LAU) and</li> <li>Nomenclature of Territorial Units for Statistics (NUTS)</li> </ul> <p><a href="https://w3id.org/launuts">https://w3id.org/launuts</a></p>
WikiCausal Corpus for Evaluation of Causal Knowledge Graph Construction
<p>Documentation on the data format and how it can be used can be found on: <a href="https://github.com/IBM/wikicausal">https://github.com/IBM/wikicausal</a> as well as our paper:</p> <pre><code>@unpublished{, author = {Oktie Hassanzadeh and Mark Feblowitz}, title = {{WikiCausal}: Corpus and Evaluation Framework for Causal Knowledge Graph Construction}, year = {2023}, doi = {10.5281/zenodo.7897996} }</code></pre> <pre>Corpus derived from Wikipedia and Wikidata. Refer to Wikipedia and Wikidata <a href="https://en.wikipedia.org/wiki/Wikipedia:Copyrights">license and terms of use</a> for more details:</pre> <ul> <li><strong>Permission is granted</strong> to copy, distribute and/or modify Wikipedia's text under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported License and, <em>unless otherwise noted</em>, the GNU Free Documentation License, unversioned, with no invariant sections, front-cover texts, or back-cover texts.</li> <li>A copy of the Creative Commons Attribution-ShareAlike 3.0 Unported License is included in the section entitled "<a href="https://en.wikipedia.org/wiki/Wikipedia:Text_of_Creative_Commons_Attribution-ShareAlike_3.0_Unported_License">Wikipedia:Text of Creative Commons Attribution-ShareAlike 3.0 Unported License</a>"</li> <li>A copy of the GNU Free Documentation License is included in the section entitled "<a href="https://en.wikipedia.org/wiki/Wikipedia:Text_of_the_GNU_Free_Documentation_License">GNU Free Documentation License</a>".</li> <li>Content on Wikipedia is covered by <a href="https://en.wikipedia.org/wiki/Wikipedia:General_disclaimer">disclaimers</a>.</li> </ul> <pre>THIS DATA IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.