Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
88
datasets available to search
ShareScore release 0.9.0
Dataset results
88 results for “Domain Analysis”
Analysis code and quantification for publication "The stress-sensing domain of activated IRE1α forms helical filaments in narrow ER membrane tubes"
<p>Analysis code and input/output files for all quantifications performed for publication entitled "The stress-sensing domain of activated IRE1α forms helical filaments in narrow ER membrane tubes." </p> <p>All questions on the analyses or code can be directed to han@walterlab.ucsf.edu</p>
Posterior samples for "Frequency-Domain Analysis of Black-Hole Ringdowns"
<p>Posterior files associated with <em>Frequency-Domain Analysis of Black-Hole Ringdowns</em> (<a href="http://arxiv.org/abs/2108.09344">arxiv:2108.09344</a>, <a href="https://journals.aps.org/prd/abstract/10.1103/PhysRevD.104.123034">Phys. Rev. D <strong>104</strong>, 123034</a>).</p> <p>Folder naming:</p> <ul> <li>{event name} <ul> <li>SNR{injection signal-to-noise ratio} <ul> <li>frequency_domain <ul> <li>{number of wavelets included in the model}W{QNM content described by lmn indices} <ul> <li>{ringdown start time in ms}_{ringdown start time prior width in ms}</li> </ul> </li> </ul> </li> <li>time_domain <ul> <li>{QNM content described by lmn indices} <ul> <li>{ringdown start time in ms}_{ringdown start time prior width in ms}</li> </ul> </li> </ul> </li> </ul> </li> </ul> </li> </ul> <p> </p> <p>For example, the file </p> <p>GW190521/SNR15/frequency_domain/1W220/12-7_01/posterior_samples.dat</p> <p>contains the results of a frequency-domain analysis of the GW190521-like injection (at signal-to-noise ratio 15) using <span class="math-tex">\(W=1\)</span> wavelet with just the fundamental <span class="math-tex">\(\ell = m = 2\)</span>, <span class="math-tex">\(n=0\)</span> QNM. It has a Gaussian prior on the ringdown start time centered 12.7 ms after the time of peak strain, with a width of 1 ms.</p> <p>The special-case two-interferometer analysis with Virgo excluded has "_HL" appended to its folder name.</p> <p>Each directory contains a 'posterior_samples.dat' file (see the 'plot_example.ipynb' for more information on reading the posterior files) and a 'sampler_output.json' which contains the log-evidence and the estimated error on the log-evidence.</p>
An analysis of the impact of domain debt in open-source projects
<p>Datasets related to the thesis paper:</p> <ul> <li><em>original_boa.txt</em> contains the intial output given by <a href="https://boa.cs.iastate.edu/boa/index.php">BOA</a>, while <em>filtered_boa.txt</em> comprises the set of repositories with at most five domain entity packages</li> <li><em>original_seart.txt</em> included data as returned by <a href="https://seart-ghs.si.usi.ch/">SEART</a></li> <li><em>database_structure.sql</em> is the SQL file that may be exploited to perform further analysis on the two datasets</li> <li><em>repos.json</em> contains the final list of repositories, as chosen in the study</li> </ul>
Data from: Molecular evolutionary analysis of nematode Zona Pellucida (ZP) modules reveals disulfide-bond reshuffling and standalone ZP-C domains
Open the record for dataset details and reuse information.
PAN19 Authorship Analysis: Cross-Domain Authorship Attribution
<p>Authorship attribution is an important problem in information retrieval and computational linguistics but also in applied areas such as law and journalism where knowing the author of a document (such as a ransom note) may enable e.g. law enforcement to save lives. The most common framework for testing candidate algorithms is the closed-set attribution task: given a sample of reference documents from a restricted and finite set of candidate authors, the task is to determine the most likely author of a previously unseen document of unknown authorship. This task may be quite challenging in <strong>cross-domain conditions</strong>, when documents of known and unknown authorship come from different domains (e.g., thematic area, genre). In addition, it is often more realistic to assume that the true author of a disputed document is not necessarily included in the list of candidates.</p> <p><strong>Fanfiction</strong> refers to fictional forms of literature which are nowadays produced by admirers ('fans') of a certain author (e.g. J.K. Rowling), novel ('Pride and Prejudice'), TV series (Sherlock Holmes), etc. The fans heavily borrow from the original work's theme, atmosphere, style, characters, story world etc. to produce new fictional literature, i.e. the so-called <strong>fanfics</strong>. This is why fanfiction is also known as transformative literature and has generated a number of controversies in recent years related to the intellectual rights property of the original authors (cf. plagiarism). Fanfiction, however, is typically produced by fans without any explicit commercial goals. The publication of fanfics typically happens online, on informal community platforms that are dedicated to making such literature accessible to a wider audience (e.g. <a href="https://www.fanfiction.net/">fanfiction.net</a>). The original work of art or genre is typically refered to as a <strong>fandom</strong>.</p> <p>This edition of PAN focuses on cross-domain attribution in fanfiction, a task that can be more accurately described as <strong>cross-fandom attribution in fanfiction</strong>. In more detail, all documents of unknown authorship are fanfics of the same fandom (target fandom) while the documents of known authorship by the candidate authors are fanfics of several fandoms (other than the target-fandom). In contrast to the PAN-2018 edition of this task, we focus on <strong>open-set attribution</strong> conditions, namely the true author of a text in the target domain is not necessarily included in the list of candidate authors.</p> <p>Each problem consists of a set of known fanfics by each candidate author and a set of unknown fanfics located in separate folders. The file <code>problem-info.json</code> that can be found in the main folder of each problem, shows the name of folder of unknown documents and the list of names of candidate author folders.</p> <p>The fanfics of known authorship belong to several fandoms (excluding the target fandom). The file <code>fandom-info.json</code> (it can be found in the main folder of each problem) provides information about the fandom of all fanfics of known authorsihp, as follows.</p> <p>The true author of each unknown document can be seen in the file <code>ground-truth.json</code>, also found in the main folder of each problem. Note that all unknown documents that are not written by any of the candidate authors belong to the <code><UNK></code> class.</p> <p>In addition, to handle a collection of such problems, the file <code>collection-info.json</code> includes all relevant information. In more detail, for each problem it lists its main folder, the language (either <code>"en"</code>, <code>"fr"</code>, <code>"it"</code>, or <code>"sp"</code>), and the encoding (always <code>UTF-8</code>) of documents.</p>
A Comprehensive Analysis of Weakly-Supervised Semantic Segmentation in Different Image Domains
<p><strong>Content</strong></p> <p>This repository contains pre-trained computer vision models, data labels, and images used in the pre-print publication "A Comprehensive Analysis of Weakly-Supervised Semantic Segmentation in Different Image Domains":</p> <ol> <li><em>ADPdevkit</em>: a folder containing the 50 validation ("tuning") set and 50 evaluation ("segtest") set of images from the Atlas of Digital Pathology database formatted in the VOC2012 style--the full database of 17,668 images is available for download from the original website</li> <li><em>VOCdevkit</em>: a folder containing the relevant files for the PASCAL VOC2012 Segmentation dataset, with both the trainaug and test sets</li> <li><em>DGdevkit</em>: a folder containing the 803 test images of the DeepGlobe Land Cover challenge dataset formatted in the VOC2012 style</li> <li><em>cues</em>: a folder containing the pre-generated weak cues for ADP, VOC2012, and DeepGlobe datasets, as required for the SEC and DSRG methods</li> <li><em>models_cnn</em>: a folder containing the pre-trained CNN models</li> <li><em>models_wsss</em>: a folder containing the pre-trained SEC, DSRG, and IRNet models, along with dense CRF settings</li> </ol> <p><strong>More information</strong></p> <p>For more information, please refer to the following article. <strong>Please cite this article when using the data set.</strong></p> <p>@misc{chan2019comprehensive,<br> title={A Comprehensive Analysis of Weakly-Supervised Semantic Segmentation in Different Image Domains},<br> author={Lyndon Chan and Mahdi S. Hosseini and Konstantinos N. Plataniotis},<br> year={2019},<br> eprint={1912.11186},<br> archivePrefix={arXiv},<br> primaryClass={cs.CV}<br> }</p> <p>For the full code released on GitHub, please visit the repository at: <a href="https://github.com/lyndonchan/wsss-analysis">https://github.com/lyndonchan/wsss-analysis</a></p> <p><strong>Contact</strong></p> <p>For questions, please contact:<br> Lyndon Chan<br> lyndon.chan@mail.utoronto.ca<br> http://orcid.org/0000-0002-1185-7961</p>
Pressure monitoring dataset and frequency domain analysis in Padua water distribution network
<p>The dataset includes:</p><ul><li>the acquired pressure signal at measurement section P036 with a high sampling frequency (indicated by "fre") for 24 hours;</li><li>the frequency domain analysis of three pressure datasets within a certain range, "omega", of frequencies</li></ul>
Spectral Domain Optical Coherence Tomography Analysis for Glaucoma
ClinicalTrials.gov study NCT01612416. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Supplementary material 3 from: Padilla-Sanchez V (2021) SARS-CoV-2 Structural Analysis of Receptor Binding Domain New Variants from United Kingdom and South Africa. Research Ideas and Outcomes 7: e62936. https://doi.org/10.3897/rio.7.e62936
United Kingdom variant interactions
Supplementary material 1 from: Padilla-Sanchez V (2021) SARS-CoV-2 Structural Analysis of Receptor Binding Domain New Variants from United Kingdom and South Africa. Research Ideas and Outcomes 7: e62936. https://doi.org/10.3897/rio.7.e62936
RBD-ACE2 complexes
Supplementary material 2 from: Padilla-Sanchez V (2021) SARS-CoV-2 Structural Analysis of Receptor Binding Domain New Variants from United Kingdom and South Africa. Research Ideas and Outcomes 7: e62936. https://doi.org/10.3897/rio.7.e62936
RBD-ACE2 interface detail
Figure 1 from: Padilla-Sanchez V (2021) SARS-CoV-2 Structural Analysis of Receptor Binding Domain New Variants from United Kingdom and South Africa. Research Ideas and Outcomes 7: e62936. https://doi.org/10.3897/rio.7.e62936
Figure 1 SARS-CoV-2 viral infection at atomic resolution. Counting eight viruses, each of which has spikes (big protrusions) and E membrane proteins (small protrusions) rainbow colored and a core in sienna color, this picture shows how the viruses approach the cell membrane (green). The ACE2 receptors are colored magenta. The field of view is 1 micrometer.
Figure 3 from: Padilla-Sanchez V (2021) SARS-CoV-2 Structural Analysis of Receptor Binding Domain New Variants from United Kingdom and South Africa. Research Ideas and Outcomes 7: e62936. https://doi.org/10.3897/rio.7.e62936
Figure 3 Detail of interface between ACE2 and SARS-CoV-2 RBD. Amino acids are labeled as well as distances. For more details please see the movies in supplementary files. Wild type is beige, UK variant is light blue and SA variant is pink (Suppl. material 2, Suppl. material 3).
Figure 2 from: Padilla-Sanchez V (2021) SARS-CoV-2 Structural Analysis of Receptor Binding Domain New Variants from United Kingdom and South Africa. Research Ideas and Outcomes 7: e62936. https://doi.org/10.3897/rio.7.e62936
Figure 2 Spike glycoprotein bound to ACE2 receptor. PDB 7DF4 (Xu et al. 2020) where ACE2 is cyan and the spike has been colored red, yellow and blue for each subunit of the trimer. This structure has been recently determined at atomic resolution. In spheres, we can see the mutations in the spike glycoprotein from the United Kingdom variant but the only mutation in the receptor binding domain (magenta) is N501Y which is labeled.
Simulation Circuits for A Time-Domain Approach to Electrical Impedance Tomography using Numerical Analysis of the Step Response
<p>SImulation circuits for the paper of the same name presented in the conference.</p>
Genome wide analysis of transcriptome modification after re-expression of PBRM1 WT and its BAH domain mutant in PBRM1 CRISPR-KO HEK293T cells.
GEO Series GSE149874. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.
Identification and functional analysis of yap1b: a new Yap family member in teleosts that evolved a divergent transcriptional activation domain.
GEO Series GSE120531. Danio rerio; Oryzias latipes. 6 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Analysis of genetically diverse macrophages reveals local and domain-wide mechanisms that control transcription factor binding and function
GEO Series GSE109965. Mus musculus. 294 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing; Methylation profiling by high throughput sequencing; Other.
PGC methylome analysis revealed large, germline-specific mouse hypomethylated DNA domains with unique genomic and epigenomic features [DNA methylation_NimbleGen]
GEO Series GSE39794. Mus musculus. 32 samples. Type: Methylation profiling by genome tiling array.
Genome-wide gene expression analysis on tibialis anterior muscle from nebulin SH3 domain deleted (Neb∆SH3) mice
GEO Series GSE47801. Mus musculus. 6 samples. Type: Expression profiling by array.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.