Abstract
In this work, we present Multiformer, a novel approach to depth-aware video panoptic segmentation (DVPS) based on the mask transformer paradigm. Our method learns ob-ject representations that are shared across segmentation, monocular depth estimation, and object tracking subtasks. In contrast to recent unified approaches that progressively refine a common object representation, we propose a hy-brid method using task-specific branches within each de-coder block, ultimately fusing them into a shared repre-sentation at the block interfaces. Extensive experiments on the Cityscapes-DVPS and SemKITTI-DVPS datasets demonstrate that Multiformer achieves state-of-the-art per-formance across all DVPS metrics, outperforming previ-ous methods by substantial margins. With a ResNet-50 backbone, Multiformer surpasses the previous best result by 3.0 DVPQ points while also improving depth estimation accuracy. Using a Swin-B backbone, Multiformer further improves performance by 4.0 DVPQ points. Multiformer also provides valuable insights into the design of multi-task decoder architectures.
| Original language | English |
|---|---|
| Title of host publication | 2025 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2025 |
| Publisher | Institute of Electrical and Electronics Engineers |
| Pages | 3301-3309 |
| Number of pages | 9 |
| ISBN (Electronic) | 979-8-3315-1083-1 |
| DOIs | |
| Publication status | Published - 8 Apr 2025 |
| Event | 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) - Tucson, United States Duration: 26 Feb 2025 → 6 Mar 2025 |
Conference
| Conference | 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) |
|---|---|
| Abbreviated title | WACV 2025 |
| Country/Territory | United States |
| City | Tucson |
| Period | 26/02/25 → 6/03/25 |
Funding
The author expresses gratitude to Dr. G. Dubbelman, Dr. D. De Geus, Prof. P.H.N. De With, and Dr. F. van der Sommen for their support and assistance. This publication is part of the NEON project with file number 17628 of the Crossover research program, which is (partly) financed by the Dutch Research Council (NWO). This work used the Dutch national compute infrastructure with the support of the SURF Cooperative using grant number EINF-5438.
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- computer vision
- monocular depth estimation
- multi-task learning
- object tracking
- scene understanding
- segmentation
- transformer networks
Fingerprint
Dive into the research topics of 'Balancing Shared and Task-Specific Representations: A Hybrid Approach to Depth-Aware Video Panoptic Segmentation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver