Skip to main navigation Skip to search Skip to main content

Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation

Research output: Chapter in Book/Report/Conference proceedingConference contributionAcademicpeer-review

129 Downloads (Pure)

Abstract

Achieving robust generalization across diverse data domains remains a significant challenge in computer vision. This challenge is important in safety-critical applications, where deep-neural-network-based systems must perform reliably under various environmental conditions not seen during training. Our study investigates whether the generalization capabilities of Vision Foundation Models (VFMs) and Unsupervised Domain Adaptation (UDA) methods for the semantic segmentation task are complementary. Results show that combining VFMs with UDA has two main benefits: (a) it allows for better UDA performance while maintaining the out-of-distribution performance of VFMs, and (b) it makes certain time-consuming UDA components redundant, thus enabling significant inference speedups. Specifically, with equivalent model sizes, the resulting VFM-UDA method achieves an 8.4× speed increase over the prior non-VFM state of the art, while also improving performance by +1.2 mIoU in the UDA setting and by +6.1 mIoU in terms of out-of-distribution generalization. Moreover, when we use a VFM with 3.6× more parameters, the VFM-UDA approach maintains a 3.3× speed up, while improving the UDA performance by +3.1 mIoU and the out-of-distribution performance by +10.3 mIoU. These results underscore the significant benefits of combining VFMs with UDA, setting new standards and baselines for Unsupervised Domain Adaptation in semantic segmentation. The implementation is available at https://github.com/tue-mps/vfm-uda.
Original languageEnglish
Title of host publication2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, CVPRW 2024
PublisherInstitute of Electrical and Electronics Engineers
Pages1172-1180
Number of pages9
ISBN (Electronic)979-8-3503-6547-4
DOIs
Publication statusPublished - 27 Sept 2024
Event2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 - Seattle, United States
Duration: 17 Jun 202421 Jun 2024

Conference

Conference2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024
Abbreviated titleCVPRW 2024
Country/TerritoryUnited States
CitySeattle
Period17/06/2421/06/24

Keywords

  • Domain adaptation
  • Computer vision
  • Semantic segmentation
  • foundation model
  • generalization
  • unsupervised domain adaptation
  • semantic segmentation
  • vision foundation model

Fingerprint

Dive into the research topics of 'Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation'. Together they form a unique fingerprint.

Cite this