
Soichiro Okazaki
Research & Development Group
Hitachi, Ltd.
Why detection confidence matters in real-world AI
What happens when an AI system correctly detects an object but assigns it a confidence score too low to trust—or incorrectly gives a high score to a false detection? In industrial environments, such errors can lead to missed inspections, unnecessary interventions, or reduced confidence in AI-assisted operations. This blog discusses research that addresses a practical challenge in deploying visual AI: how to improve recognition reliability without rebuilding or retraining the underlying detection system.
Real-world environments add another layer of complexity. New, rare, or previously unseen objects can appear long after an AI system has been deployed. Object detection, a visual AI technology that identifies and localizes objects in images, is typically trained using a fixed set of categories. As a result, conventional detectors may struggle when encountering such objects in practice.
To address this limitation, researchers have developed open-vocabulary object detection (OVOD), which allows object categories to be specified using natural-language names rather than a fixed set of predefined classes. This enables AI systems to recognize a much wider range of objects, including categories not seen during training. Recent open-vocabulary and grounding-based object detection models, such as GLIP [2], Grounding DINO [3], MM-Grounding DINO [4], and LLMDet [5] have significantly expanded the ability to detect categories beyond those seen during training. However, an important challenge remains: the confidence score—how confident the model is that a detection is correct—is not always reliable. For example, a correct rare object may receive a low score, while a similar-looking false detection may receive a high one.
To address this issue, we propose DetRefiner [1], a lightweight, model-agnostic module, that improves these confidence scores after detection. An important aspect of this approach is that existing detection systems can remain unchanged. DetRefiner works as a post-processing module, refining confidence scores without modifying the detector itself. It does this by evaluating each detection result from two perspectives: whether the category makes sense in the overall scene, and whether the object inside the bounding box visually matches that category.

Figure 1 illustrates the effect of DetRefiner on detection results. After refinement, correct objects detections such as forks, knives, paintings, and polo shirts receive higher confidence scores, making them easier to distinguish from less reliable detections. In this example, average precision (AP), a standard measure of detection accuracy, improves from 19.9 to 30.0 for rare categories and from 27.4 to 34.1 overall. These results suggest that score refinement is particularly valuable for long-tail categories with limited training examples.
How can we improve confidence reliability without retraining? The idea behind DetRefiner
How can confidence scores be improved without modifying the original detector? The key idea behind DetRefiner is to leave the detector’s predictions unchanged and focus solely on evaluating how trustworthy each detection is, by incorporating additional visual information from both the overall scene and the detected region itself.

Figure 2 shows how DetRefiner works. Starting from the candidate detections produced by an existing OVOD model, DetRefiner extracts additional visual information from the image using a frozen DINOv3 [8] encoder. “Frozen” means that DINOv3 itself is not retrained. Rather than retraining the detector, DetRefiner processes these features with a compact two-layer Transformer and produces two additional scores, which are then combined with the original detection score.
DetRefiner uses three types of DINOv3 features: a global token that summarizes the whole scene, register tokens that provide intermediate visual information, and patch tokens that represent small local regions of the image. A class vector is used to reason about the image as a whole, while patch vectors are used for local reasoning. For each predicted bounding box, the ROI Align (region of interest align) operation gathers the patch features inside that box to create a region-level representation.
During training, the class vector learns which categories are likely to appear in the scene, while the region-level vector learns whether a specific category is visually supported inside each bounding box. DetRefiner also uses MobileCLIP-based distillation [7], using knowledge from a vision-language model to improve generalization to rare and previously unseen categories.
At inference time, DetRefiner calculates an image-level score (Pcls) and a box-level score (Proi). These are combined with the original detector score (Pdet).
Pfinal = wd Pdet + wc Pcls + wp Proi
In the main experiments, the weights are wd = 0.8, wc= 0.1, and wp= 0.1. In other words, the original detector remains the primary contributor to the final score, while DetRefiner provides modest adjustments based on scene-level and box-level evidence. Because retraining is unnecessary and internal features are not required, DetRefiner can be added as a plug-and-play refinement module across different OVOD models.
Does it improve detection performance?
The next question is whether these additional confidence signals translate into measurable improvements in practice.
Table 1 shows that DetRefiner consistently improves overall detection performance across all evaluated OVOD models [6, 9], with particularly large gains for rare categories. These results show that confidence refinement can improve performance across both lightweight and high-performance detectors without altering their underlying architecture.

Why does it work? Understanding the refinement process
Figure 3 shows how the two refinement signals work together. The class-vector branch asks, in effect, “Is this category plausible somewhere in this scene?” The patch-vector branch asks, “Does the region inside this bounding box actually resemble that category?” When both signals support the base detector’s prediction, confidence increases. When they disagree, the unreliable detection score can be reduced.

What does this mean for real-world systems? Deployment considerations
From a deployment perspective, the modular design of DetRefiner allows confidence scores to be refined without modifying the underlying detector. Existing detectors, including black-box detection APIs, can remain unchanged while benefiting from improved confidence estimations.
One limitation is that DetRefiner cannot recover objects that were never proposed by the original base detector. In practice, however, this limitation can be often mitigated by using a lower detection threshold and allowing DetRefiner to recover low-confidence candidates.
Looking ahead towards more reliable open-world recognition
DetRefiner is a lightweight, detector-agnostic framework for refining open-vocabulary object detection results. It combines information about the whole scene with evidence from each detected region, improving confidence scores without modifying or retraining the base detector.
Experiments across multiple detectors show consistent accuracy gains, especially for rare categories. The results suggest that refining confidence after detection is a practical way to make open-world visual recognition more reliable while keeping existing detection systems in place.
This approach also fits naturally with Lumada, where AI is combined with existing OT, IT, and products in real-world environments. In industrial sites, mobility infrastructure, energy facilities, and retail settings, the objects encountered by visual AI can change over time. DetRefiner can improve recognition reliability while reusing existing detection assets and deployment pipelines, helping turn field image data into actionable insights.
For readers interested in technical details, please check the original paper. [1].
References
[1] S. Okazaki, T. Sasaki, and H. Ohashi, “DetRefiner: Model-Agnostic Detection Refinement with Feature Fusion Transformer,” CVPR2026 Findings, 2026.
[2] L.H. Li et al., “Grounded Language-Image Pre-training,” CVPR, 2022.
[3] S. Liu et al., “Grounding DINO: Marrying DINO with Grounded Pre-training for Open-Set Object Detection,” ECCV, 2024.
[4] X. Zhao et al., “An Open and Comprehensive Pipeline for Unified Object Grounding and Detection,” arXiv:2401.02361, 2024.
[5] S. Fu et al., “LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models,” CVPR, 2025.
[6] A. Gupta, P. Dollár, and R. Girshick, “LVIS: A Dataset for Large Vocabulary Instance Segmentation,” CVPR, 2019.
[7] P.K.A. Vasu, H. Pouransari, F. Faghri, R. Vemulapalli, and O. Tuzel, “MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training,” CVPR, 2024.
[8] O. Siméoni, et al., “DINOv3,” arXiv:2508.10104, 2025.
[9] X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui, “Open-vocabulary Object Detection via Vision and Language Knowledge Distillation,” ICLR, 2022.






