Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,063

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,063 results for “Search”

Learn how ShareScore rates datasets ↗
zenodo36/100

Search strategies for domiciliary ventilation of spinal cord injured adults

<p>The dataset includes the complete, reproducible search strategies for all literature databases searched during this project.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Search Strategies for Gene Expression Profiling Tests

<p>The dataset includes the complete, reproducible search strategies for all literature databases searched during this project. The Endnote file contain all citations considered for inclusion in the review.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Internet Search Data on Precocious Puberty 2017 to 2021

<p>We assessed in Google Trends searches for 21 precocious puberty (PP)-related terms in English internationally, in the years 2017-2021. Additionally, we assessed local searches for selected terms, in English and local languages, in countries where a rise in PP has been reported. Searches were collected in Relative Search Volumes format.</p>

opencc-by-4.0Oct 2022View details →
dryad36/100

Bird predation on Roseau cane scale as revealed by a web image search and querying a citizen monitoring database

<p>NA</p>

opencc-zeroOct 2022View details →
zenodo36/100

Supplementary Material for Nested Segmentation of Web Search Queries

<p>1. Query test of SGCL12&nbsp;[SGCL12QueryTestSet.txt]</p> <p>2. Outputs of 16 nesting algorithms on the query&nbsp;test set of SGCL12 [Input flat segmention: Saha Roy et al., SIGIR&nbsp;2012] [Nested Segmentation Outputs SGCL12.zip]</p> <p>3. Query set of TREC-WT [TREC-WTQueryTestSet.txt]</p> <p>4. Outputs of 16 nesting algorithms on the query test set of TREC-WT&nbsp;&nbsp;[Input flat segmention: Saha Roy et al.,&nbsp;SIGIR&nbsp;2012] [Nested Segmentation Outputs TREC-WT.zip]</p> <p>5.&nbsp;Code and executables&nbsp;for generating the nested segmentations for a set of queries [Code and execs for generating nested segmentations.zip]</p> <p>6.&nbsp;Code and executables&nbsp;for&nbsp;IR-based&nbsp;evaluation of a nested segmentation [Code and execs for IR evaluation of a nested segmentation.zip]</p> <p>7.&nbsp;Readme.txt</p>

opencc-by-4.0Jan 2018View details →
zenodo36/100

Head-motion and eye-gaze behavior reveal audio-visual target search strategies - dataset

<p>Participants were tasked with finding a target stimulus in the presence of a number of auditory distractors.&nbsp;</p> <p>The target stimulus (audio-only, audio-visual or visual-only) was presented at the central position for 3 seconds, after which the stimulus was moved to one of 24 positions around the participant. The other 23 positions all contained a visual distractor and 0, 1, 2, 3, 5, 7 or 11 auditory distractors (evenly spaced).&nbsp;</p> <p>Based on the eye and headtracking data, we calculated the FOV and target localization time and the maximum headrotation into the wrong direction.&nbsp;</p> <p>FOV localization time: time it took to bring target within FOV.&nbsp;<br>Target localization time: From FOV to response.</p> <p>These datasets contain both the tracking data and the summarized data.&nbsp;</p> <p>Columns "av_stim_x_y" list for each of the 24 AV stimuli (which both served as the targets and distractors), the id, the angle at which it was present during the trial and the visual and audio status.</p> <p>FUNDING:&nbsp;</p> <p>The research was supported by the Centre for Applied Hearing<br>research (CAHR) through a research consortium agreement with<br>GN Resound, Oticon, and Widex. The funders had no role in<br>study design, data collection and analysis, decision to publish, or<br>preparation of the article.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Exploring the Interaction of Code Coverage and Non-Coverage Objectives in Search-Based Test Generation

<p>Data Package for "Exploring the Interaction of Code Coverage and Non-Coverage Objectives in Search-Based Test Generation"</p> <p>This package contains data generated as part of our experiments on EvoSuite and Defects4J.</p> <p>This paper is currently under submission.</p> <p>This package contains experimental data and the algorithms that generated them in folder "framework\test\". Specifically, in the folder "framework\test\Experiments", the data folders contain the suites generated by each technique for each project and fault ID.&nbsp;</p> <p>If you have questions, please contact Afonso Fontes at afonsohfontes@gmail.com.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Replication package for: Search Complementarities, Aggregate Fluctuations, and Fiscal Policy

<p>Fern&aacute;ndez-Villaverde J, Mandelman F, Yu Y, Zanetti F. Search complementarities, aggregate fluctuations, and fiscal policy.&nbsp;<em>Review of Economic Studies</em></p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

The digital appendices for (2017) Robert Fortune - Plant Hunter From the Scottish Borders to the Pacific Rim in Search of Plants by David Kay Ferguson. ISBN: 9787566414519

<p>At the back of the 2017 book "Robert Fortune - Plant Hunter From the Scottish Borders to the Pacific Rim in Search of Plants" by David Kay Ferguson. ISBN: 9787566414519 there is a CD with two files on it.&nbsp; Notes &amp; References. This is an archive of these Microsoft Word files.</p> <p>This archiving is done without the authors permission (I did try) and is intended to allow access to the data held within.</p>

opencc-by-4.0May 2024View details →
dryad36/100

Explaining dimorphism polymorphism: Stronger interspecific sexual differences may be favored when females search for mates in the presence of congeners

<p>Why are some species sexually dimorphic while other closely related species are not? While all females in genus <em>Strauzia</em> share a multiply-banded wing pattern typical of many other true fruit flies, males of four species have noticeably elongated wings with banding patterns "coalesced" into a continuous dark streak across much of the wing. We take an integrative phylogenetic approach to explore the evolution of this dimorphism and develop general hypotheses underlying the evolution of wing dimorphism in flies. We find that the origin of coalesced and other darkened male wing patterns correlate with the inferred origin of host plant sharing in <em>Strauzia.</em> While wing shape among non-host-sharing species tended to be conserved across the phylogeny, shapes of male wings for <em>Strauzia</em> species sharing the same host plant were more different from one another than expected under Brownian models of evolution and overall rates of wing shape change differed between non-host-sharing species and host-sharing species. A survey of North American Tephritidae finds just three other genera with specialist species that share host plants. Host-sharing species in these genera also have wing patterns unusual for each genus. Only genus <em>Eutreta </em>is like<em> Strauzia </em>in<em> </em>having the unusual wing patterns only in males, and of genera that have multiple species sharing hosts, only in <em>Eutreta </em>and <em>Strauzia</em> do males hold territories while females search for mates. We hypothesize that in species that share host plants, those where females actively search for males<em> </em>in the presence of congeners may be more likely to evolve sexually dimorphic wing patterns.</p>

opencc-zeroMay 2024View details →
zenodo36/100

The IBEM Dataset: a large printed scientific image dataset for indexing and searching mathematical expressions

<p>The IBEM dataset consists of 600 documents with a total number of 8272 pages, containing 29603 isolated and 137089 embedded Mathematical Expressions (MEs). The objective of the IBEM dataset is to facilitate the indexing and searching of MEs in massive collections of STEM documents. The dataset was built by parsing the LaTeX source files of documents from the <a href="https://www.cs.cornell.edu/projects/kddcup/datasets.html">KDD Cup Collection</a>. Several experiments can be carried out with the IBEM dataset ground-truth (GT): ME detection and extraction, ME recognition, etc.</p> <p>&nbsp;</p> <p>The dataset consists of the following files:</p> <ul> <li>&ldquo;IBEM.json&rdquo;: file containing the IBEM GT information. The data is firstly organized by pages, then by the type of expression (&ldquo;embedded&rdquo; or &ldquo;displayed&rdquo;), and lastly by the GT of each individual ME. For each ME we provide: <ul> <li>xy page-level coordinates, reported as relative (%) to the width/height of the page image.</li> <li>&ldquo;split&rdquo; attribute indicating the number of fragments in which the ME has been split. MEs can be split over various lines, columns or pages. The LaTeX transcript of split MEs have been exactly replicated (entire LaTeX definition) for each fragment.</li> <li>&ldquo;latex&rdquo; original transcript as extracted from the LaTeX source files of the documents. This definition can contain user-defined macros. In order to be able to compile these expressions, each page includes the preamble of the source files containing the defined macros and the packages used by the authors of the documents.</li> <li>&ldquo;latex_expand&rdquo; transcript reconstructed from the output stream of the LuaLaTeX engine in which user-defined macros have been expanded. The transcript has the same visual representation as the original transcript, with the addition that the LaTeX definitions are tokenized, the order of sub/super script elements have been fixed, and matrices have been transformed to arrays.</li> <li>&ldquo;latex_norm&rdquo; transcript resulting from applying an extra normalization process to the &ldquo;latex_expand&rdquo; expression. This normalization process includes removing font information such as slant, style, and weight.</li> </ul> </li> <li>&ldquo;partitions/*.lst&rdquo;: files containing list of pages forming the partition sets.</li> <li>&ldquo;pages/*.jpg&rdquo;: individual pages extracted from the documents.</li> </ul> <p>The dataset is partitioned into various sets as provided for the ICDAR 2021 Competition on Mathematical Formula Detection. The ground-truth related to this competition, which is included in this dataset version, can also be found <a href="https://zenodo.org/record/4757865">here</a>. More information about the competition can be found in the following paper:</p> <p>D. Anitei, J.A. S&aacute;nchez, J.M. Fuentes, R. Paredes, and J.M. Bened&iacute;. ICDAR 2021 Competition on Mathematical Formula Detection. In ICDAR, pages 783&ndash;795, 2021.</p> <p>&nbsp;</p> <p>For ME recognition tasks, we recommend rendering the &ldquo;latex_expand&rdquo; version of the formulae in order to create standalone expressions that have the same visual representation as MEs found in the original documents (see attached python script &ldquo;extract_GT.py&rdquo;). Extracting MEs from the documents based on coordinates is more complex, as special care is needed to concatenate the fragments of split expressions. Baseline results for ME recognition tasks will soon be made available.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Adaptive Parameter Control for Search-Based Unit Test Generation — Replication Package

<h1>Running the experiments</h1> <p>Prerequisites for running the experiments:</p> <ul> <li><a href="https://www.docker.com/" target="_blank" rel="noopener">Docker</a></li> <li><a href="https://python-poetry.org/" target="_blank" rel="noopener">Poetry</a></li> </ul> <p>Steps to run the experiments:</p> <ol> <li>Download the experiment zip file you wish to run (<code>single_parameter_experiment.zip</code> or <code>multi_parameter_experiment.zip</code>).</li> <li>Un-zip the file.</li> <li>Open a terminal and navigate to the unzipped folder (e.g. <code>cd single_parameter_experiment</code>).</li> <li>Run <code>poetry install --only main</code> to install all dependencies.</li> <li>To run the experiment, run <code>poetry run python run_experiment.py</code>.</li> <li>All results can be found in the folder <code>data/</code>.</li> </ol> <p>The modules used for the experiment are defined in the file <code>experiment_modules.py</code> and to see the experiment configuration, look in <code>run_experiment.py</code>.</p> <p><strong>Warning</strong>: the experiments take several weeks to run on a single machine, therefore it is advisable to split the experiments based on modules and run them in parallel.</p> <h1>Running the analysis</h1> <p>Prerequisites for running the analysis:</p> <ul> <li><a href="https://conda.io/projects/conda/en/latest/user-guide/getting-started.html" target="_blank" rel="noopener">Conda</a></li> </ul> <p>Steps to run the analysis:</p> <ol> <li>Download the analysis zip file (<code>analysis-adaptive-parameter-control.zip</code>).</li> <li>Un-zip the file.</li> <li>Open a terminal and navigate to the unzipped folder (e.g. <code>cd analysis-adaptive-parameter-control</code>).</li> <li>Run the following command to install the conda environment and all dependencies: <code>conda env create -f environment.yml</code></li> <li>If you want to re-run the Bayesian models locally on your machine, follow the <strong>optional</strong> step below, otherwise download and unzip the trace data from the replication package,&nbsp;i.e., <code>Trace data single.zip</code> and <code>Trace data multi.zip</code>.&nbsp;</li> <li>Place the <code>.nc</code> files in the corresponding folder: <code>analysis-adaptive-parameter-control/single_parameter/</code> or <code>analysis-adaptive-parameter-control/multi_parameter/</code>. <ol> <li>E.g. the&nbsp;<code>coverage_rate_model_single_parameter.nc</code> goes in the <code>single_parameter</code> folder, while the&nbsp; <code>coverage_rate_model_multi_parameter.nc</code> goes in the&nbsp;<code>multi_parameter</code> folder.</li> </ol> </li> <li>Navigate to the notebooks folder (<code>Notebooks/</code>).</li> <li>Open a notebook of choice (<code>coverage_rate_multi_parameter.ipynb</code>, <code>coverage_rate_single_parameter.ipynb</code>, <code>final_coverage_multi_parameter.ipynb</code>, <code>final_coverage_single_parameter.ipynb</code>, <code>overhead_model_multi_parameter.ipynb</code>, or <code>overhead_model_single_parameter.ipynb</code>).</li> <li>Navigate to the section called "Data analysis" and run all cells in order.</li> </ol> <h3>(Optional) Running the Bayesian models locally before the analysis.</h3> <ol> <li>Navigate to the notebooks folder (<code>Notebooks/</code>).</li> <li>Open a notebook of choice (<code>coverage_rate_multi_parameter.ipynb</code>, <code>coverage_rate_single_parameter.ipynb</code>, <code>final_coverage_multi_parameter.ipynb</code>, <code>final_coverage_single_parameter.ipynb</code>, <code>overhead_model_multi_parameter.ipynb</code>, or <code>overhead_model_single_parameter.ipynb</code>).</li> <li>Navigate to the section called "Model specification" and run the three notebook cells.</li> </ol> <p><strong>Warning</strong>: this will take a long time, if you don't have the time, use the following alternative instead</p> <h1>Data</h1> <p>The data from when we ran the experiments is available in the <code>Single data.zip</code> and <code>Multi data.zip</code> files.</p> <p>The structure of these are the following:</p> <ul> <li>There are folders for each module the experiment was run on, further divided into each unique run. All these folders include:&nbsp; <ul> <li>Coverage reports.</li> <li>Complete logs for the unique run.</li> <li>A timeline over controlled parameter values during the test generation process.</li> <li>The complete Pynguin configuration for the run.</li> <li>The generated test suite.</li> </ul> </li> <li>There is one <code>statistics.csv</code> file containing some information about each run and their branch coverage timelines.</li> </ul> <h1>Running the parameter assignment analysis</h1> <p>Prerequisites for running the parameter assignment analysis:</p> <ul> <li><a href="https://conda.io/projects/conda/en/latest/user-guide/getting-started.html" target="_blank" rel="noopener">Conda</a></li> </ul> <p>Steps to run the parameter assignment analysis:</p> <ol> <li>Download the analysis zip file (<code>parameter-assignment.zip</code>).</li> <li>Un-zip the file.</li> <li>Open a terminal and navigate to the unzipped folder (e.g. <code>cd parameter-assignment</code>).</li> <li>Run the following command to install the conda environment and all dependencies: <code>conda env create -f environment.yml</code></li> <li>Navigate to the notebooks folder (<code>Notebooks/</code>).</li> <li>Open the notebook&nbsp;<code>parameter_assignment_analysis.ipynb</code>.</li> <li>Run all cells in order.</li> </ol>

opencc-by-4.0May 2024View details →
zenodo36/100

An infrared search for R Coronae Borealis Stars

<p>J-band lightcurves, medium resolution near-IR spectra, spectroscopic classifications and lightcurve and color-based priorities of all sources described in the paper An infrared census of R Coronae Borealis Stars II.&nbsp;</p>

opencc-by-4.0Jul 2024View details →
dryad36/100

Data from: The influence of the number of tree searches on maximum likelihood inference in phylogenomics

<p>Maximum likelihood (ML) phylogenetic inference is widely used in phylogenomics. As heuristic searches most likely find suboptimal trees, it is recommended to conduct multiple (e.g., ten) tree searches in phylogenetic analyses. However, beyond its positive role, how and to what extent multiple tree searches aid ML phylogenetic inference remains poorly explored. Here, we found that a random starting tree was not as effective as the BioNJ and parsimony starting trees in inferring ML gene tree and that RAxML-NG and PhyML were less sensitive to different starting trees than IQ-TREE. We then examined the effect of the number of tree searches on ML tree inference with IQ-TREE and RAxML-NG, by running 100 tree searches on 19,414 gene alignments from 15 animal, plant, and fungal phylogenomic datasets. We found that the number of tree searches substantially impacted the recovery of the best-of-100 ML gene tree topology among 100 searches for a given ML program. In addition, all of the concatenation-based trees were topologically identical if the number of tree searches was ≥ 10. Quartet-based ASTRAL trees inferred from 1 to 80 tree searches differed topologically from those inferred from 100 tree searches for 6 /15 phylogenomic datasets. Lastly, our simulations showed that gene alignments with lower difficulty scores had a higher chance of finding the best-of-100 gene tree topology and were more likely to yield the correct trees.</p>

opencc-zeroJul 2024View details →
zenodo36/100

Supplemental Data for "Identifying novel variants of small molecules through database search of mass spectra"

<p>Supplemental Dataset 1 contains GNPS dataset information for the large scale search of GNPS vs Pubchem.</p> <p>Supplemental Dataset 2a contains the top scoring exact mode hit for each spectrum against PubChem and COCONUT. Files are split into "*chunk*" files of up to 10 million records each.</p> <p>Supplemental Dataset 2b contains the top scoring variable mode hit for each spectrum against COCONUT.</p> <p>Supplemental Dataset 3 contains mass spectra provided by Waters Corporation to analyze impurities of Imatinib.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Search-Based Test Data Generation for SQL Queries: Appendix

<p>The appendix of our ICSE 2018 paper &quot;Search-Based Test Data Generation for SQL Queries: Appendix&quot;.</p> <p>The appendix contains:</p> <ul> <li>The queries from the three open source systems we used in the evaluation of our tool (the industry software system is not part of this appendix, due to privacy reasons)</li> <li>The results of our evaluation.</li> <li>The source code of the tool. Most recent version can be found at&nbsp;https://github.com/SERG-Delft/evosql.</li> <li>The results of the tuning procedure we conducted before running the final evaluation.</li> </ul>

opencc-by-4.0Feb 2018View details →
zenodo36/100

SubDiv17: A Dataset for Investigating Subjectivity in the Visual Diversification of Image Search Results

<p>This dataset facilitates the comparison of approaches aiming at the diversification of image search results. The dataset was explicitly designed for general-purpose, multi-topic queries and provides multiple ground truth annotations to allow for the exploration of the subjectivity aspect in the general task of diversification. The dataset provides images and their metadata retrieved from Flickr for around 200 complex queries. Additionally, to encourage experimentations (and cooperations) from different communities such as information and multimedia retrieval, a broad range of pre-computed descriptors is provided. The dataset was successfully validated during the MediaEval 2017 Retrieving Diverse Social Images task using 29 submitted runs. For more information, please see&nbsp;<a href="https://doi.org/10.1145/3204949.3208122">https://doi.org/10.1145/3204949.3208122</a>.</p>

opencc-by-4.0Jun 2018View details →
zenodo36/100

A Search for a Surviving White Dwarf Companion in SN~1006 - photometric dataset

<p>This is a photometry catalogue obtained by DECam to find hot surviving WD companions to the SN1006 supernova. Please acknowledge the science paper Kerzendorf et al. 2017 if you use this dataset.&nbsp;</p>

opencc-by-4.0Aug 2017View details →
zenodo36/100

Search for huntingtin interactors in online databases – 2018/08/08

<p><strong>Project</strong>&nbsp;- Huntingtin structure-function open lab notebook.&nbsp;</p> <p><strong>Rationale</strong> - To identify different huntingtin interaction partners.&nbsp;</p> <p><strong>Overview</strong> - Different online databases which detail protein interaction partners were searched for huntingtin protein interaction partners.&nbsp;Data detailing huntingtin interaction partners from 9 different databases was extracted and simplified &ndash; worksheets 1-15.&nbsp;&nbsp;The information from each database was collated &ndash; worksheet 16.&nbsp;Huntingtin protein interaction partners were ranked according to the number of databases they were found in as well as the number of different experiments detailing the interaction with huntingtin &ndash; worksheet 17.&nbsp;</p>

opencc-by-4.0Aug 2018View details →
zenodo36/100

Comparative analysis of huntingtin interaction partner database searches with Dr. Maiuri's work - 2018/08/31

<p><strong>Project&nbsp;</strong>- Huntingtin structure-function open lab notebook.&nbsp;</p> <p><strong>Rationale</strong>&nbsp;- To compare the ROS-specific huntingtin interaction partners identified by Dr. Tamara Maiuri with those detailed in existing databases.&nbsp;</p> <p><strong>Overview</strong>&nbsp;- Previously, different online databases which detail protein interaction partners were searched for huntingtin protein interaction partners.&nbsp;Following completion of this initial analysis, Dr. Tamara Maiuri posted in her open notebook a detailed list of high and medium confidence ROS-specific huntingtin interacting proteins:&nbsp;<a href="https://zenodo.org/record/1319540">https://zenodo.org/record/1319540</a>. A comparison of these interactors with those identified in the&nbsp;previously mined databases is briefly detailed.&nbsp;</p>

opencc-by-4.0Aug 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record