When to Adapt: Multi-Signal Domain Shift Detection for Efficient Training-Free Adaptation in Open-Vocabulary Segmentation
Michele Antonazzi*, Alejandra C. Hernandez*, José Araujo, Olov Andersson, Patric Jensfelt
* Equal contribution
Michele Antonazzi*, Alejandra C. Hernandez*, José Araujo, Olov Andersson, Patric Jensfelt
* Equal contribution
Robust and reliable perception is essential for autonomous robots operating in real-world environments, particularly in long-term missions where environmental conditions may change significantly over time. Although recent advances in Visual Foundation Models (VFMs) have improved open-vocabulary semantic segmentation, these models can still suffer from domain shift, which can significantly degrade performance if they are not adapted to the current environment. Training-free domain adaptation is a relevant paradigm for adaptation, consisting of adjusting the model online using lightweight adapters. Recent approaches apply this on a per-frame basis, which is impractical for deployments on resource-constrained robotic hardware. To tackle this, we propose a multi-signal domain shift detection method for training-free continual test-time adaptation (TF-CTTA) in open-vocabulary segmentation. Our method leverages temporal coherence across consecutive frames by monitoring and combining complementary aspects of domain shift (visual change, adapter mismatch, and semantic drift) to trigger adaptation only when needed. We validate our approach on a benchmark including indoor and outdoor environments and using real robotic data. We demonstrate that our approach maintains segmentation accuracy while substantially reducing adaptations, making training-free adaptation practical and feasible for long-term, real-world robotic deployments.
Robots performing long-term missions, such as last-mile delivery or mobile assistance, move across multiple and changing environments: indoor and outdoor scenarios, new buildings, and varying illumination and weather. This variability generates domain shift, which degrades the performance of open-vocabulary segmentation models if they are not adapted to the current environment. Training-free adaptation methods such as SemLA adjust the model online by retrieving and merging low-rank adapters from a library, without any backpropagation. However, they perform this operation at every frame, introducing a substantial runtime overhead on resource-constrained robotic hardware. Since the images perceived by a robot are highly correlated in time, we instead detect domain shifts in the image stream and trigger adaptation only when needed.
Our domain shift detector monitors three complementary signals (visual change, adapter mismatch, and semantic drift) to decide when it is worth swapping the adapters loaded into the segmentation model.
Our framework builds on SemLA, a library \(\mathcal{L}\) of LoRA adapters, each indexed by a descriptor \(C_s\) computed as the average CLIP embedding of its training domain. When adaptation is triggered, the top-\(K\) adapters closest to the current domain are merged and attached to the open-vocabulary segmentation model. Our domain shift detector decides when to adapt by monitoring three signals that capture complementary aspects of domain shift:
The signals are combined with a hierarchical fusion. Since \(S_1\) is highly sensitive, when it fires a confirmation from \(S_2\) or \(S_3\) is required; when there is no abrupt visual change, both \(S_2\) and \(S_3\) must agree. The signals operate on the CLIP embeddings and predictions already available in the pipeline, so they add no model inference.
Overview of the framework for training-free continual test-time adaptation, integrating SemLA and our multi-signal domain shift detector.
We design a continual test-time adaptation benchmark tailored to robotics, with sequences of 16 domains alternating indoor (Scannet, Scannet++) and outdoor (ACDC, MUSES) scenes. The model is always tested on unseen domains, as the adapter trained on the current domain is discarded. Results are averaged over 50 randomly sampled sequences. Our method stays within 1.4 mIoU of per-frame adaptation while triggering adaptation only 0.8% of the time, and outperforms the other domain shift detection strategies.
| Method | mIoU ↑ | NA ↓ |
|---|---|---|
| Zero-shot | 38.6 | 0 |
| Entropy | 38.4 | 0.3 |
| DSS | 41.8 | 0.5 |
| Periodic | 43.9 | 4.3 |
| Multi-signal (ours) | 44.8 | 3.9 |
| Per-frame (SemLA) | 46.2 | 518 |
Mean segmentation accuracy (mIoU) and number of adaptations per domain (NA) across sequences of 16 indoor and outdoor domains. Per-frame adaptation represents the accuracy upper bound.
Individually, the signals are sensitive: \(S_1\) and \(S_3\) trigger frequently, while \(S_2\) barely fires, as adapter mismatch builds gradually. Combined, they are jointly selective, discarding spurious triggers while preserving detection coverage. The full system achieves the best trade-off between accuracy and number of adaptations.
Trade-off between segmentation accuracy and number of adaptations for the individual signals and their combinations, averaged over 50 sequences of 8 domains each.
We profile the pipeline on an NVIDIA Jetson AGX Orin mounted on a Spot robot. Adaptation is the bottleneck: each adapter swap takes 550–700ms, more than CLIP and segmentation inference combined in most configurations. Triggering adaptation only when needed raises the throughput up to 3.2× compared to per-frame adaptation.
| Power mode | Models | Domain nav. | Seg. model | Adapt. cost | FPS (per-frame) | FPS (ours) |
|---|---|---|---|---|---|---|
| High (60W) | Large | 150ms | 230ms | 550ms | 1.1 | 2.6 |
| Small | 90ms | 140ms | 550ms | 1.3 | 4.2 | |
| Low (30W) | Large | 350ms | 840ms | 700ms | 0.5 | 0.8 |
| Small | 110ms | 490ms | 700ms | 0.7 | 1.7 |
Per-component latency and throughput of per-frame adaptation and our triggered adaptation (NA = 4) on a Jetson AGX Orin.
We deploy our method on a real-world sequence of 2735 frames captured with an Odin sensor while navigating indoor and outdoor areas of the KTH campus. The visual change signal fires 42 times, but the hierarchical fusion confirms only 12 adaptations, rejecting 71% of the activations as false positives. The confirmed adaptations correspond to visually distinct domain transitions: corridor changes, indoor-to-outdoor passages, and room type shifts. On the Jetson, these 12 adaptations add only 7s of overhead, compared to 25 minutes under per-frame adaptation.
Our domain shift detector on a real-world indoor/outdoor sequence. Each image pair covers consecutive adaptation events: the first frame is where adaptation is triggered, the second is from the stable period before the next trigger.
@article{antonazzi2026whentoadapt,
title = {When to Adapt: Multi-Signal Domain Shift Detection for Efficient Training-Free Adaptation in Open-Vocabulary Segmentation},
author = {Antonazzi, Michele and Hernandez, Alejandra C. and Araujo, José and Andersson, Olov and Jensfelt, Patric},
year = {2026},
}