Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,037
datasets available to search
ShareScore release 0.9.0
Dataset results
1,037 results for “large-scale”
Figure 5 in Large-scale snake genome analyses provide insights into vertebrate development
Figure 5. Evolution of blind and infrared-sensitive snakes (A) Specialized traits of blind snakes. (B) Gene regression underlies eye degeneration in blind snakes. Lost genes are shown in green. Heatmap showing the significant down-regulation of the expression of genes in blind snake eyes (QN, quantile normalized). (C) BSD-CNE-associated genes enriched in odontogenesis and bone-development-related GO terms. (D) Amino acids sequences alignment (left) and enzymatic activity (right) of CHIA. The amino acids in red are positively selected sites in blind snakes. The error bar represents mean ± SD of enzymatic activity. The enzymatic activity is significantly higher in Diard's blind snake (Student's t-test p value = 8.78e-09). (E) Infrared sensing-related genes in infrared-sensitive snakes. PSGs, REGs, and variation-shared genes are indicated in purple, blue, and red, respectively. Average expression levels (represented by FPKM) of these genes in seven tissues of the brown-spotted pitviper are shown. See also Figure S5 and Tables S4 and S5.
Figure 4 in Large-scale snake genome analyses provide insights into vertebrate development
Figure 4. Genome features associated with snake-specific sense organ evolution, highlighting PSGs, REGs, newly evolved genes, SD-CNEassociated genes, lost genes, and SSSV-associated genes (A) Lost and highly expressed genes in snake eyes. Diagram of the snake eye showing lost genes in blue. Heatmap of highly expressed genes (QN, quantile normalized). (B) Diagram of the snake inner ear. Twelve SD-CNE-associated genes and four SSSV-associated genes are involved in ear development. Expression levels (scaled fragments per kilobase per million mapped reads [FPKM]) of sound perception-related PSGs and REGs in 10 keeled slug snake (Pber) tissues are shown in circles. Circle size is positively correlated to the expression level. (C) Diagram of the taste transduction process. Solid lines indicate direct interactions, and dotted lines indicate indirect interactions. See also Table S3.
Figure 1 in Large-scale snake genome analyses provide insights into vertebrate development
Figure 1. Phylogeny of snakes Maximum-likelihood phylogenetic tree inferred from whole-genome sequences of 31 species. Divergence times of all nodes were estimated by r8s with whole-genome sequences using six calibration points (Figures S1I and S1J). All genomes generated in this study are in red. Maps were taken from those in a previous study.24
Figure S5 in Large-scale snake genome analyses provide insights into vertebrate development
Figure S5. Genome evolution of snake sensory system, related to Figures 4 and 5 (A) Expression patterns of genes associated with eye development in human embryos at different embryonic stages (4.7–8 weeks post conception [wpc]) and red cornsnake embryos at 10, 30, and 50 days post oviposition (dpo). (B) KMT2C specific amino acids residue lost in snakes and its associated snakes divergent CNE (SD-CNE). The black line portrays gene structure with blue blocks representing CDSs and the numbers indicate the exact positions on the European glass lizard reference genome. Three segments of snake-specific amino acid residue loss are manifested by internal numbers ''4002–4004,'' ''2110–2112,'' and ''1983–1987.'' One SD-CNE located at 30 regulate region of KMT2C is marked by a red block. (C) Alignment of SSSV that inserted into the 5 kb upstream of PDZD7 transcription start site. This SSSV conveys a new EBF1 binding site with a regulatory potential (RP) score: 0.78. (D) Taste transduction involved genes were expressed in tongue and brain of three snakes (plumbeous water snake [Hplu], keeled slug snake [Pber], and Asian vine snake [Apra]). The numbers in cells indicate the scaled FPKM values. (E) Comparison of eye structures between Diard's blind snake and Asian vine snake, keeled slug snake. The magnifications were marked as ''*×'' at the bottomright corner of each histological sections. (F) GOChord plot (produced by GOplot package) of blind snake divergent CNEs involved coding gene-enriched GO terms (p value <0.05). Left half of GOChord displayed genes of different GO terms and the right showed the GO term descriptions. Each gene was linked to a GO term by the colored bands. (G) REVIGO clusters of significantly overrepresented (p value <0.01) GO terms for blind snake REGs. (H) Expression levels of TRPV4 and TRPA1 in seven tissues of brown spotted pitviper. (I) Genome locations of pitviper diverged conserved non-coding elements (PVD-CNEs), and the infrared-sensitive python and boa divergent conserved non-coding elements (ISPBD-CNEs) in PMP22 and NFIB.
Dataset for large-scale self-organisation in dry turbulent atmospheres
<p>Datasets for all figures in "Large-scale self-organisation in dry turbulent atmospheres". The data comes from a simulation of the Boussinesq equations in a triply periodic domain of vertical height H and horizontal dimension L = 32H, in the presence of gravity, a stable mean density gradient, and solid body rotation in the vertical direction. Datasets are in TXT format except for two dimensional spectra, which are stored in NetCDF format.</p>
Fig. 2. A in A large-scale study on the seroprevalence of Toxoplasma gondii infection in humans in Iran
Fig. 2. A GIS map of IgM seroprevalence of Toxoplasma gondii (Nicolle et Manceaux, 1908) in different provinces of Iran, during 2015–2020 (ND – no data).
Fig. 3. A in A large-scale study on the seroprevalence of Toxoplasma gondii infection in humans in Iran
Fig. 3. A GIS map of IgG seroprevalence of Toxoplasma gondii (Nicolle et Manceaux, 1908) in different provinces of Iran, during 2015–2020.
Fig. 1 in A large-scale study on the seroprevalence of Toxoplasma gondii infection in humans in Iran
Fig. 1. The seroprevalence plot of Toxoplasma gondii (Nicolle et Manceaux, 1908) in Iran in different age groups (m-month; y-year).
A Large-Scale Dataset of 4G, NB-IoT, and 5G Non-Standalone Network Measurements
<p>Mobile networks have become highly complex systems. In order to better understand how network features affect performance and suggest additional improvements, it is crucial to examine them from an empirical perspective. In the following, we present a large-scale dataset of measurements collected over fourth generation (4G) and fifth generation (5G) operational networks, providing Long Term Evolution (LTE), Narrowband Internet of Things (NB-IoT) and 5G New Radio (NR) connectivity. We collected our dataset during a period of seven weeks in Rome, Italy, by performing several tests on the infrastructures of two major mobile network operators (MNOs). The open-sourced dataset has enabled multi-faceted analyses of network deployment, coverage, and end-user performance, and can be further used for designing and testing artificial intelligence (AI) and machine learning (ML) solutions for network optimization tasks.</p> <p><br>If you use our dataset in your research, we kindly request that you cite the following paper:</p> <p>K. Kousias <em>et al</em>., "A Large-Scale Dataset of 4G, NB-IoT, and 5G Non-Standalone Network Measurements," in <em>IEEE Communications Magazine</em>, vol. 62, no. 5, pp. 44-49, May 2024, doi: 10.1109/MCOM.011.2200707.</p>
Domain-adaptive Data Synthesis for Large-scale Supermarket Product Recognition
<p><strong>Domain-Adaptive Data Synthesis for Large-Scale Supermarket Product Recognition</strong></p> <p>This repository contains the data synthesis pipeline and synthetic product recognition datasets proposed in [1].</p> <p><strong>Data Synthesis Pipeline:</strong></p> <p>We provide the Blender 3.1 project files and Python source code of our data synthesis pipeline <em>pipeline.zip, </em>accompanied by the<em> </em><a href="https://github.com/taesungp/contrastive-unpaired-translation">FastCUT</a> models used for synthetic-to-real domain translation<em> models.zip</em>. For the synthesis of new shelf images, a product assortment list and product images must be provided in the corresponding directories <em>products/assortment/</em> and <em>products/img/</em>. The pipeline expects product images to follow the naming convention <em>c</em>.png, with <em>c</em> corresponding to a GTIN or generic class label (e.g., 9120050882171.png). The assortment list, <em>assortment.csv</em>, is expected to use the sample format [<em>c, w, d, h</em>], with <em>c</em> being the class label and <em>w, d,</em> and <em>h</em> being the packaging dimensions of the given product in mm (e.g., [4004218143128, 140, 70, 160]). The assortment list to use and the number of images to generate can be specified in <em>generateImages.py </em>(see comments). The rendering process is initiated by either executing <em>load.py</em> from within Blender or within a command-line terminal as a background process. </p> <p><strong>Datasets:</strong></p> <ul> <li><strong>SG3k</strong> - Synthetic GroZi-3.2k (SG3k) dataset, consisting of 10,000 synthetic shelf images with 851,801 instances of 3,234 GroZi-3.2k products. Instance-level bounding boxes and generic class labels are provided for all product instances.</li> <li><strong>SG3kt</strong> - Domain-translated version of SGI3k, utilizing GroZi-3.2k as the target domain. Instance-level bounding boxes and generic class labels are provided for all product instances.</li> <li><strong>SGI3k</strong> - Synthetic GroZi-3.2k (SG3k) dataset, consisting of 10,000 synthetic shelf images with 838,696 instances of 1,063 GroZi-3.2k products. Instance-level bounding boxes and generic class labels are provided for all product instances.</li> <li><strong>SGI3kt</strong> - Domain-translated version of SGI3k, utilizing GroZi-3.2k as the target domain. Instance-level bounding boxes and generic class labels are provided for all product instances.</li> <li><strong>SPS8k</strong> - Synthetic Product Shelves 8k (SPS8k) dataset, comprised of 16,224 synthetic shelf images with 1,981,967 instances of 8,112 supermarket products. Instance-level bounding boxes and GTIN class labels are provided for all product instances.</li> <li><strong>SPS8kt</strong> - Domain-translated version of SPS8k, utilizing SKU110k as the target domain. Instance-level bounding boxes and GTIN class labels for all product instances.</li> </ul> <p>Table 1: Dataset characteristics. </p> <table> <tbody> <tr> <td><strong>Dataset</strong></td> <td><strong>#images</strong></td> <td><strong>#products</strong></td> <td><strong>#instances</strong></td> <td> <strong>labels </strong></td> <td><strong>translation</strong></td> </tr> <tr> <td>SG3k</td> <td>10,000</td> <td>3,234</td> <td>851,801</td> <td>bounding box & generic class¹</td> <td>none</td> </tr> <tr> <td>SG3kt</td> <td>10,000</td> <td>3,234</td> <td>851,801</td> <td>bounding box & generic class¹</td> <td>GroZi-3.2k</td> </tr> <tr> <td>SGI3k</td> <td>10,000</td> <td>1,063</td> <td>838,696</td> <td>bounding box & generic class²</td> <td>none</td> </tr> <tr> <td>SGI3kt</td> <td>10,000</td> <td>1,063</td> <td>838,696</td> <td>bounding box & generic class²</td> <td>GroZi-3.2k</td> </tr> <tr> <td>SPS8k</td> <td>16,224</td> <td>8,112</td> <td>1,981,967</td> <td>bounding box & GTIN</td> <td>none</td> </tr> <tr> <td>SPS8kt</td> <td>16,224</td> <td>8,112</td> <td>1,981,967</td> <td>bounding box & GTIN</td> <td>SKU110k</td> </tr> </tbody> </table> <p> </p> <p><strong>Sample Format</strong></p> <p>A sample consists of an RGB image (i.png) and an accompanying label file (i.txt), which contains the labels for all product instances present in the image. Labels use the YOLO format [c, x, y, w, h].</p> <p>¹SG3k and SG3kt use generic pseudo-GTIN class labels, created by combining the GroZi-3.2k food product category number <em>i</em> (1-27) with the product image index <em>j </em>(j.jpg)<em>, </em>following the convention<em> i0000j </em>(e.g., 13000097).</p> <p>²SGI3k and SGI3kt use the generic GroZi-3.2k class labels from <a href="https://arxiv.org/abs/2003.06800">https://arxiv.org/abs/2003.06800</a>.</p> <p><strong>Download and Use</strong><br>This data may be used for non-commercial research purposes only. If you publish material based on this data, we request that you include a reference to our paper [1].</p> <p>[1] Strohmayer, Julian, and Martin Kampel. "Domain-Adaptive Data Synthesis for Large-Scale Supermarket Product Recognition." <em>International Conference on Computer Analysis of Images and Patterns</em>. Cham: Springer Nature Switzerland, 2023.</p> <p>BibTeX citation:</p> <pre>@inproceedings{strohmayer2023domain, title={Domain-Adaptive Data Synthesis for Large-Scale Supermarket Product Recognition}, author={Strohmayer, Julian and Kampel, Martin}, booktitle={International Conference on Computer Analysis of Images and Patterns}, pages={239--250}, year={2023}, organization={Springer} }</pre>
Data & figures: Comparison between Large-Scale Observed and Simulated Antarctic Sea-Ice Variability Response to Changes in Atmospheric and Oceanic Circulation
<p>These are the model data, key figures, and Python code generated during the project titled “Comparison between Large-Scale Observed and Simulated Antarctic Sea-Ice Variability Response to Changes in Atmospheric and Oceanic Circulation." This project was undertaken during a 3-month research scholarship at the Alfred Wegener Institute Helmholtz Centre for Polar and Marine Research, funded by the Helmholtz Visiting Researcher Grant, a program promoted by the Helmholtz Information and Data Science Academy (HIDA). Statistical methods pertain to the coupling of sea surface temperature and Antarctic sea-ice interactions. These methods can be applied to observations, reanalysis, and earth system model data</p>
JusBrasilRec: A large-scale dataset of user sessions for recommendations on the legal domain
<p><strong>JusBrasilRec: A large-scale dataset of user sessions for recommendations on the legal domain</strong></p> <p>The proliferation of legal documents in various formats and their dispersion across multiple courts present a significant challenge for users seeking precise matches to their information requirements. Despite notable advancements in legal information retrieval systems, research into legal recommender systems remains limited. A plausible factor contributing to this scarcity could be the absence of extensive publicly accessible datasets or benchmarks.</p> <p>Jusbrasil (<a href="https://www.jusbrasil.com.br">https://www.jusbrasil.com.br</a>) is known as the largest legal search portal in Brazil. It provides an online environment where users can find the legal documents that best match their information needs. With millions of user interactions to billions of documents containing different artifacts related to law in Brazil, Jusbrasil appears as a large-scale test bed for advancing research on the still scarce area of legal recommender systems. </p> <p>Therefore, we collected and made available the <strong>JusBrasilRec</strong>, a dataset containing user sessions from Jusbrasil for recommendations on the legal domain. Additionally, we also computed and made available a TF-IDF matrix from the textual content of the documents in Jusbrasil. The following files are available for download from JusBrasilRec:</p> <ul> <li><strong>jusbrasilrec_dataset.zip:</strong> a compacted file containing the user sessions;</li> <li><strong>jusbrasilrec_tfidf_matrix.zip:</strong> a compacted file containing the TF-IDF matrix;</li> <li><strong>readme.txt:</strong> a text file explaining the content and format of the previous files.</li> </ul> <p><strong>How to cite the dataset:</strong> Marcos Aurélio Domingues, Edleno Silva de Moura, Leandro Balby Marinho and Altigran da Silva. A Large Scale Benchmark for Session-based Recommendations on the Legal Domain. Artificial Intelligence and Law. 2023.</p>
Required data for simulating a typical large-scale urban traffic network
Open the record for dataset details and reuse information.
Input data to model multiple effects of large-scale deployment of grass in crop-rotations at European scale
Open the record for dataset details and reuse information.
Data for: Characterization of large-scale preferential flow across continental United States
Open the record for dataset details and reuse information.
Data from: Large-scale eDNA sampling and hierarchical modeling elucidates the importance of stream habitat for eastern hellbender (<em>Cryptobranchus a. alleganiensis</em>) occupancy and eDNA detection
Open the record for dataset details and reuse information.
Increasing sustainability in palaeoproteomics by optimizing digestion times for large-scale archaeological bone analyses
Open the record for dataset details and reuse information.
Large-scale neural recordings with single neuron resolution using Neuropixels probes in human cortex
Open the record for dataset details and reuse information.
Data for: Large-scale long-term passive-acoustic monitoring reveals spatiotemporal activity patterns of boreal bats
Open the record for dataset details and reuse information.
Data from: Using large-scale tropical dry forest restoration to test successional theory
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.