Skip to main navigation Skip to search Skip to main content

Balancing Shared and Task-Specific Representations: A Hybrid Approach to Depth-Aware Video Panoptic Segmentation

Research output: Chapter in Book/Report/Conference proceedingConference contributionAcademicpeer-review

32 Downloads (Pure)

Abstract

In this work, we present Multiformer, a novel approach to depth-aware video panoptic segmentation (DVPS) based on the mask transformer paradigm. Our method learns ob-ject representations that are shared across segmentation, monocular depth estimation, and object tracking subtasks. In contrast to recent unified approaches that progressively refine a common object representation, we propose a hy-brid method using task-specific branches within each de-coder block, ultimately fusing them into a shared repre-sentation at the block interfaces. Extensive experiments on the Cityscapes-DVPS and SemKITTI-DVPS datasets demonstrate that Multiformer achieves state-of-the-art per-formance across all DVPS metrics, outperforming previ-ous methods by substantial margins. With a ResNet-50 backbone, Multiformer surpasses the previous best result by 3.0 DVPQ points while also improving depth estimation accuracy. Using a Swin-B backbone, Multiformer further improves performance by 4.0 DVPQ points. Multiformer also provides valuable insights into the design of multi-task decoder architectures.
Original languageEnglish
Title of host publication2025 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2025
PublisherInstitute of Electrical and Electronics Engineers
Pages3301-3309
Number of pages9
ISBN (Electronic)979-8-3315-1083-1
DOIs
Publication statusPublished - 8 Apr 2025
Event2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) - Tucson, United States
Duration: 26 Feb 20256 Mar 2025

Conference

Conference2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
Abbreviated titleWACV 2025
Country/TerritoryUnited States
CityTucson
Period26/02/256/03/25

Funding

The author expresses gratitude to Dr. G. Dubbelman, Dr. D. De Geus, Prof. P.H.N. De With, and Dr. F. van der Sommen for their support and assistance. This publication is part of the NEON project with file number 17628 of the Crossover research program, which is (partly) financed by the Dutch Research Council (NWO). This work used the Dutch national compute infrastructure with the support of the SURF Cooperative using grant number EINF-5438.

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • computer vision
  • monocular depth estimation
  • multi-task learning
  • object tracking
  • scene understanding
  • segmentation
  • transformer networks

Fingerprint

Dive into the research topics of 'Balancing Shared and Task-Specific Representations: A Hybrid Approach to Depth-Aware Video Panoptic Segmentation'. Together they form a unique fingerprint.

Cite this