Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
131
datasets available to search
ShareScore release 0.9.0
Dataset results
131 results for “queries”
Pre-built Symphonypy reference objects that can be downloaded and used to map new query datasets
<p>Pre-built Symphonypy reference objects that can be downloaded and used to map new query datasets. The same data as in <a href="https://zenodo.org/record/5090425#.Y9EmLC8w3tg">https://zenodo.org/record/5090425</a>, but for Python port for Symphony.</p> <p>The Symphony algorithm is used to perform reference mapping to these atlases. </p> <ul> <li>Paper: <a href="https://www.nature.com/articles/s41467-021-25957-x">https://www.nature.com/articles/s41467-021-25957-x</a></li> <li>Usage: <a href="https://github.com/potulabe/symphonypy">https://github.com/potulabe/symphonypy</a></li> </ul> <p><strong>References available for download:</strong></p> <ol> <li>10x PBMCs Atlas (pbmcs_10x_reference.h5ad)</li> <li>Pancreatic Islet Cells Atlas (pancreas_plate-based_reference.h5ad)</li> <li>Fetal Liver Hematopoiesis Atlas (fetal_liver_reference_3p.h5ad)</li> <li>Healthy Fetal Kidney Atlas (kidney_healthy_fetal_reference.h5ad)</li> <li>T cell CITE-seq atlas (tbru_ref.h5ad)</li> <li>Cross-tissue Inflammatory Immune Atlas (zhang_reference.h5ad)</li> <li>Tabula Muris Senis (FACS) Atlas (TMS_facs_reference.h5ad)</li> </ol>
Housing Transitions QUERI
ClinicalTrials.gov study NCT05312229. IPD Sharing: YES. Countries: 1. Publications: 1.
Stroke Prediction Through Internet Search Queries
ClinicalTrials.gov study NCT04755959. IPD Sharing: UNDECIDED. Countries: 0. Publications: 4.
Abstractive Snippet Generation (query-biased anchor context)
<p>Abstractive Snippet Generation (query-biased anchor context)</p>
Data and Intermediate Queries for: Evaluating institutional open access performance: Methodology, challenges and assessment
<p>This package provides the publicly sharable data for two related articles, alongside the key processing steps used to generate it.</p>
WikiProject Clinical Trials snapshot of 2021 query results
<p><strong>WikiProject Clinical Trials snapshot of 2021 query results<br> December 2021<br> data results from project queries</strong></p> <p><strong>Background</strong><br> Wikidata is a community and database within the Wikipedia ecosystem.</p> <p>WikiProject Clinical Trials is a community project in Wikidata to curate data related to clinical trials for its use in the Wikipedia ecosystem or export to elsewhere. Visit the project at <a href="https://www.wikidata.org/wiki/Wikidata:WikiProject_Clinical_Trials">https://www.wikidata.org/wiki/Wikidata:WikiProject_Clinical_Trials</a></p> <p><strong>About this data</strong><br> Presented here is a collection of data snapshots which demonstrate model outputs of queries to the data which WikiProject Clinical Trials has curated in Wikidata.</p> <p>Elements of these data snapshots which are worth consideration include these listed queries, the social and cultural need which these queries seek to fulfill, and the results of these queries. Beyond what these queries actually report, persons familiar with the field of clinical research are able to look at this information and consider the origins of some of it, note what data is here that would be difficult to find elsewhere, and also to imagine what data would be useful here but which is absent.</p> <p>While these queries are representative of an early attempt of what WikiProject Clinical Trials aspires to deliver, for many variations of these queries, Wikidata's content is incomplete and there will be gaps. For example, the project acquired excellent data about the researchers, research publications, research projects, and research departments at Vanderbilt University. Consequently, these queries can profile that university relatively well. In comparison, the Wikidata has no comparable data collection for any other university, so there is incompleteness.</p> <p>We share these datasets because reproducing the state of Wikidata for past times is difficult, and wish to report these results as a milestone of our progress till now.</p> <p><strong>Data snapshots</strong></p> <p>These datasets were collected 19 December 2021. See the original queries in WikiProject Clinical Trials.</p> <ol> <li>Lists which order topics by count of clinical trials <ol> <li>List of medical conditions, ordered by count of clinical trials which feature them</li> <li>List of research interventions, ordered by count of clinical trials which feature them</li> <li>List of research institutions, ordered by count of times each served as research site</li> <li>List of people, ordered by count of times each served as principal investigator</li> <li>List of funders, ordered by count of clinical trials</li> </ol> </li> <li>Example queries <ol> <li>List of clinical trials where medical condition is Zika fever</li> <li>List of clinical trials where medical intervention is an instance of or subclass of a COVID-19 vaccine</li> <li>List of clinical trials where the place of research was Vanderbilt University or any of its research sites</li> <li>List of clinical trials where the principal investigator was Alp Ikizler</li> <li>List of clinical trials with principal investigator and their affiliation</li> <li>List of clinical trials where the principal investigator was affiliated with Vanderbilt University</li> <li>Bubble chart of organizations by number of clinical trials</li> <li>List of clinical trials where the funder was the Patient-Centered Outcomes Research Institute</li> <li>List of clinical trials where the sponsor was Pfizer</li> </ol> </li> <li>For conversation <ol> <li>Count of the genders of principal investigators</li> <li>List of clinical trials where the principal investigator was female</li> <li>Count of principal investigators by occupation</li> </ol> </li> <li>Scope of Wikidata's content <ol> <li>List of clinical trials</li> <li>Count of clinical trials</li> <li>List of clinical trial registries, ordered by count of clinical trials cataloged in Wikidata</li> </ol> </li> </ol>
Retrieving API knowledge from Tutorials and Stack Overflow based on Natural Language Queries
<p>The replication package of PLAN</p>
Query structures from Figure 4 of Comparison of Substructure Search Systems
<p>These 29 SMARTS queries come from Figure 4 on page 193 of:</p> <p> Martin G. Hicks and Clemens Jochum,<br> Substructure search systems. 1. Performance comparison of the<br> MACCS, DARC, HTSS, CAS Registry MVSSS, and S4 substructure search systems,<br> Journal of Chemical Information and Computer Sciences 1990 30 (2), 191-199<br> <a href="https://doi.org/10.1021/ci00066a018">https://doi.org/10.1021/ci00066a018</a></p> <p>The SMARTS queries are in Figure 4 (page 193). A copy of the original table is in 10.1021_ci00066a018_Figure_4.pdf . The LibreOffice spreadsheet file 10.1021_ci00066a018.ods contain two sheets. The "Overview" sheet gives a summary of the data set. The "Query structures" sheet contains the data from Figure 4.</p> <p>The Hicks and Jochum paper does not give the query as SMARTS strings. Instead, the queries were translated to SMARTS in the following paper:</p> <p> Dimitris K. Agrafiotis, Victor S. Lobanov, Maxim Shemanarev,<br> Dmitrii N. Rassokhin, Sergei Izrailev, Edward P. Jaeger,<br> Simson Alex, and Michael Farnum<br> Efficient Substructure Searching of Large Chemical Libraries:<br> The ABCD Chemical Cartridge<br> Journal of Chemical Information and Modeling 2011 51 (12), 3113-3130<br> <a href="https://doi.org/10.1021/ci200413e">https://doi.org/10.1021/ci200413e</a></p> <p>Two of the SMARTS patterns were changed to produce this dataset. During cross-checking with the original paper, and with the help of SMARTSviewer, I believe I identified two mistakes in the conversion, described in the README.</p>
Raw Dataset of DiTU - A query-based heuristic language feature usage analysis method
<p>The dataset contains raw query results and aggregated metric analysis of ~3k commits in our large-scale language feature usage analysis. Please refer to the README.md in our anonymous GitHub repository for concrete usage guidance.</p>
Navigating and Querying Answer Sets: How Hard Is It Really and Why? (Empirical Case Study)
Open the record for dataset details and reuse information.
S-Cypher: A Temporal Query Language on the Temporal Property Graph Model
Open the record for dataset details and reuse information.
Raptor@GenomeBiology: RNA-Seq queries
<table> <thead> <tr> <th scope="col">Filename</th> <th scope="col">Compressed size</th> <th scope="col">Uncompressed size</th> </tr> </thead> <tbody> <tr> <td>100.fa.zst</td> <td>67 KiB</td> <td>226 KiB</td> </tr> <tr> <td>1k.fa.zst</td> <td>623 KiB</td> <td>2.2 MiB</td> </tr> <tr> <td>10k.fa.zst</td> <td>5.5 MiB</td> <td>23 MiB</td> </tr> <tr> <td>50k.fa.zst</td> <td>28 MiB</td> <td>112 MiB</td> </tr> <tr> <td>100k.fa.zst</td> <td>55 MiB</td> <td>224 MiB</td> </tr> </tbody> </table> <p> </p> <table> <thead> <tr> <th scope="col">Filename</th> <th scope="col">SHA256 compressed</th> </tr> </thead> <tbody> <tr> <td>100.fa.zst</td> <td>9ac1db1540190c64939ac986aabffd399c8e46bd89389d8befa891b0e22db299</td> </tr> <tr> <td>1k.fa.zst</td> <td>6ad28ee01e5ced2cdbee7183e2af2bb95fde326b697b39cf88a0fd6446f9f441</td> </tr> <tr> <td>10k.fa.zst</td> <td>59901bb367923b0ae4e00264effe28e7aa90b8d86bad7e3bedd8e952e05f22b4</td> </tr> <tr> <td>50k.fa.zst</td> <td>3f57aebfa5a1d6e52095e8f7c687943604430caf000e726db9a9b6c4244d918d</td> </tr> <tr> <td>100k.fa.zst</td> <td>5d0357735a2acbf47115b7af5137b76bf5ed9073b800b529b01d6a66ade43119</td> </tr> </tbody> </table> <p> </p> <table> <thead> <tr> <th scope="col">Filename</th> <th scope="col">SHA256 uncompressed</th> </tr> </thead> <tbody> <tr> <td>100.fa.zst</td> <td>2d1babccdf197baca65aebb08a45755e19532321ec9100aad2769850ce2de55d</td> </tr> <tr> <td>1k.fa.zst</td> <td>af2857d962f38a65a6754ef2ea15b4d181b00a08fc736b07ba6fc0e8ee3086b4</td> </tr> <tr> <td>10k.fa.zst</td> <td>c014d438ea24c0e21ebc4b6188a26866cb757d425a43230a01fd6bcf1b53c1d0</td> </tr> <tr> <td>50k.fa.zst</td> <td>a21e2802c47ad7a8b056fd45990c15727f44d20e474f4af4c2c88a404221fd54</td> </tr> <tr> <td>100k.fa.zst</td> <td>5327800a5300865fdbbed2785923566d8a15145dd26ea63023ff902b2c9a827c</td> </tr> </tbody> </table>
FedUP: A Pay-as-you-go Federated SPARQL Query Engine Powered by Random Walks
<p>This repository contains all the queries and datasets used to generate the figures presented in the experimental study of the "FedUP: A Pay-as-you-go Federated SPARQL Query Engine Powered by Random Walks" paper. The data comes from the <a href="https://github.com/dice-group/LargeRDFBench">LargeRDFBench</a> and <a href="https://github.com/MaastrichtU-IDS/federatedQueryKG/blob/main/usecaseFedShop.md">FedShop</a> benchmarks. Datasets have been cleaned to be ingested in tested federated query engines. More details are available in the paper and the <a href="https://github.com/GDD-Nantes/FedUP-experiments">GitHub repository</a>.</p> <p><strong>queries.tar.gz: </strong>LargeRDFBench and FedShop queries.</p> <p><strong>datasets.tar.gz: </strong>Cleaned LargeRDFBench and FedShop datasets.</p> <p><strong>summaries.tar.gz: </strong>Pre-built summaries for tested federated query engines such as <a href="https://2014.eswc-conferences.org/sites/default/files/papers/paper_50.pdf">HiBiSCuS</a> and <a href="https://svn.aksw.org/papers/2018/SEMANTICS_CostFed/public.pdf">CostFed</a>. </p>
Implementing Cognitive Behavioral Therapy for Insomnia (SWELL): Function QUERI 3.0
ClinicalTrials.gov study NCT07216261. IPD Sharing: YES. Countries: 1. Publications: 0.
Implementing a Group Physical Therapy Program for Veterans (GroupPT): Function QUERI 3.0
ClinicalTrials.gov study NCT07215572. IPD Sharing: YES. Countries: 1. Publications: 0.
Implementing a Mobile Health Application for Women Veterans With Urinary Incontinence (MyHealtheBladder): Function QUERI 3.0
ClinicalTrials.gov study NCT07219433. IPD Sharing: YES. Countries: 1. Publications: 0.
Adult Congenital Heart Disease Registry (QuERI)
ClinicalTrials.gov study NCT01659411. IPD Sharing: Not stated. Countries: 0. Publications: 1.
A tumor multi-region query reveals novel DNA methylation targets linked to ccRCC poor outcome and metastasis
GEO Series GSE206049. Homo sapiens. 169 samples. Type: Methylation profiling by array.
Handwriting query data
<p>The upload contains 1) the results of a query with 26 respondents who assessed the consistency of 20 Early Medieval handwriting specimens and 2) the questionnaire which was used in the query. The questionnaire is translated into English, but it was originally in Finnish. The upload is related to an upcoming journal article.</p>
An Algorithm for Context-Free Path Queries Over Graph Databases
<p>Recorded presentation for the paper "An Algorithm for Context-Free Path Queries Over Graph Databases", for SBLP 2020.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.