Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
59
datasets available to search
ShareScore release 0.9.0
Dataset results
59 results for “Fine-grained”
Reconstructed Three-decade Global Fine-Grained Nighttime Light Dataset
<p>Nighttime light (NTL) is a foundational data source for studying human activities from a remote sensing perspective. This dataset from 1992 to 2021 is the first global long-term and fine-grained NTL observations. It represents a milestone in facilitating the study of human activities. It is created by a new super-resolution model DeepNTL which converts DMSP-OLS images into NPP-VIIRS images. Compared with baseline models, including RCAN, SwinIR and AutoEncoder, DeepNTL has the hightest accuracy and best generalization ability for untrained years. It provides a good extension of NPP-VIIRS to the early years, and the future annual NPP-VIIRS data can be directly appended to this dataset by users own. More information about the dataset can be found in the "read_me.txt" file. Technical details and evaluations are presented in our paper. Any questions are welcome to be sent to jinyuguo23@m.fudan.edu.cn.</p>
Animal Recognition Using Methods Of Fine-Grained Visual Analysis - YOLOv5 Breed Classification Dataset (Tsinghua Dogs)
<p>Tsinghua Dogs Dataset with ground truth labels for breeds in YOLOv5 format.</p>
The Delta Maintainability Model: Measuring Maintainability of Fine-Grained Code Changes Technical Report
<p>Technical Report and dataset for "The Delta Maintainability Model: Measuring Maintainability of Fine-Grained Code Changes" submitted to the 2nd International Conference on Technical Debt (TechDebt 2019) - Montréal, Canada - May 26–27, 2019.</p>
High-level Synthesis of Fine-grained Weakly Consistent C Concurrency
<p>This repository consists of all the experimental data of PhD thesis funded by EPSRC. </p>
A transcriptome for the early-branching fern Botrychium lunaria enables fine-grained resolution of population structure
<p>File S1: Alignment of Botrychium CRY2cA sequences used to infer the genus-level phylogeny.</p> <p>File S2: Peptide sequences used to infer the orthogroups named according species names or identifiers (see Table S1).</p> <p>File S3: Alignments of the orthogroup sequences subset used to infer the phylogenomy.</p> <p>File S4: Output files from modeltest-ng named by orthogroup names.</p> <p>File S5: Output files from raxml-ng named by orthogroup names.</p>
Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
<p>QA-Feedback used in the paper: Fine-Grained Human Feedback Gives Better Rewards for Language Model Training</p>
Data for: Pre-vegetation, single-thread rivers sustained by cohesive, fine-grained bank sediments: Mesoproterozoic Stoer Group, NW Scotland
Open the record for dataset details and reuse information.
Data from: Fine-grained neural coding of bodies and body parts in human visual cortex
Open the record for dataset details and reuse information.
Data from: Movies reveal the fine-grained organization of infant visual cortex
Open the record for dataset details and reuse information.
Fine-Grained Just-In-Time Defect Prediction - Appendix
<p>Defect prediction models focus on identifying defect-prone code elements, for example to allow practitioners to allocate testing resources on specific subsystems and to provide assistance during code reviews. While the research community has been highly active in proposing metrics and methods to predict defects on long-term periods (i.e., at release time), a recent trend is represented by the so-called short-term defect prediction (i.e., at commit-level). Indeed, this strategy represents an effective alternative in terms of effort required to inspect files likely affected by defects. Nevertheless, the granularity considered by such models might be still too coarse. Indeed, existing commit-level models highlight an entire commit as defective even in cases where only specific files actually contain defects. </p> <p>In this paper, we first investigate to what extent commits are partially defective; then, we propose a novel fine-grained just-in-time defect prediction model to predict the specific files, contained in a commit, that are defective. Finally, we evaluate our model in terms of (i) performance and (ii) the extent to which it decreases the effort required to diagnose a defect. Our study highlights that: (1) defective commits are frequently composed of a mixture of defective and non- defective files, (2) our fine-grained model can accurately predict defective files with an AUC-ROC up to 82% and (3) our model would allow practitioners to save inspection efforts with respect to standard just-in-time techniques.</p>
Fine-grained automated visual analysis of herbarium specimens for phenological data extraction: an annotated dataset of reproductive organs in Strepanthus herbarium specimens
<p>This dataset contains annotations of 31 herbarium specimens of <em>Streptanhus tortuosus Kellogg</em> for which we have we carefully and manually drew and annotated the contours of four reproductive organs: “bud”, “flower”, “immature fruit” and “mature fruit”.</p> <p>The dataset can be used to assess the ability of automated methods to count and detect precisely the shapes of these reproductive organs, with a view to conducting phenological studies.</p> <p>The annotations are formatted in accordance with the COCO data format, a usual format for object detection tasks in the field of Computer Vision. The annotations are divided into two files:</p> <ul> <li>train_21_full_masks.json contains the mask coordinates and labels of 21 herbarium sheets that can be used for training models</li> <li>test_10_full_masks.json contains the mask coordinates and labels of 10 other herbarium that can be used as a groundtruth file for evaluating the predictions, typically with the COCO evaluation scripts (<a href="https://github.com/cocodataset/cocoapi">https://github.com/cocodataset/cocoapi</a>)</li> </ul> <p>Please refer to the following publication for a first assessment of this dataset with a Mask-RCNN approach:</p> <p><em>H. Goëau, A. Mora-Fallas, J. Champ, N. Love, S. Mazer, E. Mata-Montero, A. Joly, P. Bonnet. </em>2020. New fine-grained method for automated visual analysis of herbarium specimens: a case study for phenological data extraction. <em>Applications in Plant Sciences </em></p> <p> </p> <p> </p> <p> </p> <p> </p>
Data from: Fine-grained adaptive divergence in an amphibian: genetic basis of phenotypic divergence and the role of non-random gene flow in restricting effective migration among wetlands
Adaptive ecological differentiation among sympatric populations is promoted by environmental heterogeneity, strong local selection and restricted gene flow. High gene flow, on the other hand, is expected to homogenize genetic variation among populations and therefore prevent local adaptation. Understanding how local adaptation can persist at the spatial scale at which gene flow occurs has remained an elusive goal, especially for wild vertebrate populations. Here, we explore the roles of natural selection and nonrandom gene flow (isolation by breeding time and habitat choice) in restricting effective migration among local populations and promoting generalized genetic barriers to neutral gene flow. We examined these processes in a network of 17 breeding ponds of the moor frog Rana arvalis, by combining environmental field data, a common garden experiment and data on variation in neutral microsatellite loci and in a thyroid hormone receptor (TRβ) gene putatively under selection. We illustrate the connection between genotype, phenotype and habitat variation and demonstrate that the strong differences in larval life history traits observed in the common garden experiment can result from adaptation to local pond characteristics. Remarkably, we found that haplotype variation in the TRβ gene contributes to variation in larval development time and growth rate, indicating that polymorphism in the TRβ gene is linked with the phenotypic variation among the environments. Genetic distance in neutral markers was correlated with differences in breeding time and environmental differences among the ponds, but not with geographical distance. These results demonstrate that while our study area did not exceed the scale of gene flow, ecological barriers constrained gene flow among contrasting habitats. Our results highlight the roles of strong selection and nonrandom gene flow created by phenological variation and, possibly, habitat preferences, which together maintain genetic and phenotypic divergence at a fine-grained spatial scale.
Fine-grained classification of journal articles by relying on multiple layers of information through similarity network fusion: the case of the Cambridge Journal of Economics
<p>prova</p>
Replication package for Fine-grained, accurate and scalable source differencing
<p>Replication package for the article <em><strong>Fine-grained, accurate and scalable source differencing</strong></em> accepted at the 46th International Conference on Software Engineering (ICSE 2024).</p> <p>List of files:</p> <ul> <li><strong>dataset</strong>: the folder containing the four datasets used in the paper, including the file-pairs and the script to regenerate them.</li> <li><strong>qualitative_experiment_diffs</strong>: the folder containing the GUI used by the participants of the qualitative experiment on the 100 cases.</li> <li><strong>gumtree</strong>: the folder containing the code of GumTree used to run the experiments.</li> <li><strong>analysis</strong>: the folder containing the CSV files and notebooks for the analysis of the results.</li> </ul>
A fine-grained dataset named iSOOD for sewage outfalls objective detection in natural environments
<p><strong>Basic Information:</strong></p> <p>The 10481 images in iSOOD were captured using UAVs and handheld cameras by individuals from the river basin in China. Our study has carefully annotated these images to ensure accuracy and consistency. The iSOOD has undergone technical validation utilizing the YOLOv5 objective detection model. The iSOOD have been publicly released after undergoing desensitization, with the goal of promoting interdisciplinary collaboration and accelerating advancements in the intelligence watershed management. We expect that the iSOOD dataset is anticipated to inspire further research on the SOs detection and the control of pollution migration paths, and serve as a fundamental resource for the use of advanced deep learning visual technology in environmental monitoring.</p> <p><strong>Usage Policy:</strong><br>If you plan to use our data in a scientific analysis paper, we strongly recommend contacting us in advance to seek opinions, and consider our contributions in the acknowledgments or as co-authors.</p>
Let's Trace It: Fine-Grained Serverless Benchmarking using Synchronous and Asynchronous Orchestrated Applications - Dataset
<p>This dataset contains the raw collected traces, preprocessed versions of the traces, and summary figures for the data associated with our manuscript <em>Let's Trace It: Fine-Grained Serverless Benchmarking using Synchronous and Asynchronous Orchestrated Applications.</em></p> <p>It contains over 7.5 million (7 564 830) traces of the ten applications integrated with ServiBench. The measurements were conducted on AWS Lambda in the us-east-1 region in late 2021 and early 2022. For more details on how the traces were collected we refer to our manuscript.</p> <p>For details on how to replicate our existing analysis on this dataset, we refer to https://github.com/ServiBench/ReplicationPackage</p>
Retrieving Data Constraint Implementations Using Fine-Grained Code Patterns
<p>Business rules are an important part of the requirements of software systems that are meant to support an organization. These rules describe the operations, definitions, and constraints that apply to the organization. Within the software system, business rules are often translated into constraints on the values that are required or allowed for data, called data constraints. Business rules are subject to frequent changes, which in turn require changes to the corresponding data constraints in the software. The ability to efficiently and precisely identify where data constraints are implemented in the source code is essential for performing such necessary changes.</p> <p>In this paper, we introduce Lasso, the first technique that automatically retrieves the method and line of code where a given data constraint is enforced. Lasso is based on traceability link recovery approaches and leverages results from recent research that identified line-of-code level implementation patterns for data constraints. We implement three versions of Lasso that can retrieve data constraint implementations when they are implemented with any one of 13 frequently occurring patterns. We evaluate the three versions on a set of 299 data constraints from 15 real-world Java systems, and find that they improve method-level link recovery by 30%, 70%, and 163%, in terms of true positives within the first 10 results, compared to their text-retrieval-based baseline. More importantly, the Lasso variants correctly identify the line of code implementing the constraint inside the methods for 68% of the 299 constraints.</p>
UGS-1m: Fine-grained urban green space mapping of 34 major cities in China based on the deep learning framework
<p>Urban green space (UGS) is an important component in the urban ecosystem and has great significance to the urban ecological environment. The UGS-1m product provides the fine-grained UGS maps of 34 major cities/areas in China, which is generated based on a deep learning (DL) framework.</p> <p>The DL framework consists of a generator and a discriminator. The generator is a fully convolutional network designed for UGS extraction (UGSNet), which integrates attention mechanisms to improve the discrimination to UGS, and employs a point rending strategy for edge recovery. The discriminator is a fully connected network aiming to deal with the domain shift between images. To support the model training, an urban green space dataset (UGSet) with a total number of 4,454 samples of size 512×512 is provided. Code for the UGSet and the UGSet will be soon available at: https://github.com/liumency/UGS-1m. </p> <p>The main steps to obtain UGS-1m can be summarized as follows: a) Firstly, the UGSNet will be pre-trained on the UGSet in order to get a good starting training point for the generator; b) After pre-training on the UGSet, the discriminator is responsible to adapt the pre-trained UGSNet to different cities/areas through adversarial training; c) Finally, the UGS results of the 34 major cities/areas in China (UGS-1m) are obtained using 2,343 Google Earth images with a data frame of 7'30" in longitude and 5'00" in latitude, and a spatial resolution of nearly 1.1 meters. Evaluating the performance of the proposed approach on samples from Guangzhou city shows the validity of the UGS-1m products, with an overall accuracy of 87.4% and an F1 score of 81.14%. </p>
Dataset of "Inferring Fine-grained Traceability Links between Javadoc Comments and JUnit Test Code"
<p>Dataset of "Inferring Fine-grained Traceability Links between Javadoc Comments and JUnit Test Code"</p> <p>- study object, true link, sentence, test code snippet, experiment result</p>
Animal Recognition Using Methods Of Fine-Grained Visual Analysis - Kashtanka Pets (400 Hand-labelled Images - Cats & Dogs, Single Folder)
<p>400 images (200 cats, 200 dogs) hand-labelled by Maria E. with head and body bounding box labels in YOLOv5 format. Images are in a single folder, no separate folders for cats and dogs.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.