Abstract
Monocular depth estimation and semantic segmentation are two fundamental goals of scene understanding. Due to the advantages of task interaction, many works have studied the joint-task learning algorithm. However, most existing methods fail to fully leverage the semantic labels, ignoring the provided context structures and only using them to supervise the prediction of segmentation split, which limits the performance of both tasks. In this paper, we propose a network injected with contextual information (CI-Net) to solve this problem. Specifically, we introduce a self-attention block in the encoder to generate an attention map. With supervision from the ideal attention map created by semantic label, the network is embedded with contextual information so that it could understand the scene better and utilize correlated features to make accurate prediction. Besides, a feature-sharing module (FSM) is constructed to make the task-specific features deeply fused, and a consistency loss is devised to ensure that the features mutually guided. We extensively evaluate the proposed CI-Net on NYU-Depth-v2, SUN-RGBD, and Cityscapes datasets. The experimental results validate that our proposed CI-Net could effectively improve the accuracy of semantic segmentation and depth estimation.
| Original language | English |
|---|---|
| Pages (from-to) | 18167-18186 |
| Number of pages | 20 |
| Journal | Applied Intelligence |
| Volume | 52 |
| Issue number | 15 |
| DOIs | |
| State | Published - Dec 2022 |
| Externally published | Yes |
Keywords
- Attention mechanism
- Depth estimation
- Semantic segmentation
- Task interaction
Fingerprint
Dive into the research topics of 'CI-Net: a joint depth estimation and semantic segmentation network using contextual information'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver