Skip to main navigation Skip to search Skip to main content

Learning From Architectural Redundancy: Enhanced Deep Supervision in Deep Multipath Encoder-Decoder Networks

  • Ying Luo
  • , Jinhu Lu*
  • , Xiaolong Jiang
  • , Baochang Zhang
  • *Corresponding author for this work
  • Beihang University
  • Alibaba Group Holding Ltd.

Research output: Contribution to journalArticlepeer-review

Abstract

Deep encoder-decoders are the model of choice for pixel-level estimation due to their redundant deep architectures. Yet they still suffer from the vanishing supervision information issue that affects convergence because of their overly deep architectures. In this work, we propose and theoretically derive an enhanced deep supervision (EDS) method which improves on conventional deep supervision (DS) by incorporating variance minimization into the optimization. A new structure variance loss is introduced to build a bridge between deep encoder-decoders and variance minimization, and provides a new way to minimize the variance by forcing different intermediate decoding outputs (paths) to reach an agreement. We also design a focal weighting strategy to effectively combine multiple losses in a scale-balanced way, so that the supervision information is sufficiently enforced throughout the encoder-decoders. To evaluate the proposed method on the pixel-level estimation task, a novel multipath residual encoder is proposed and extensive experiments are conducted on four challenging density estimation and crowd counting benchmarks. The experimental results demonstrate the superiority of our EDS over other paradigms, and improved estimation performance is reported using our deeply supervised encoder-decoder.

Original languageEnglish
Pages (from-to)4271-4284
Number of pages14
JournalIEEE Transactions on Neural Networks and Learning Systems
Volume33
Issue number9
DOIs
StatePublished - 1 Sep 2022

Keywords

  • Convolutional neural networks (CNN)
  • crowd counting
  • deep supervised learning
  • encoder-decoder networks

Fingerprint

Dive into the research topics of 'Learning From Architectural Redundancy: Enhanced Deep Supervision in Deep Multipath Encoder-Decoder Networks'. Together they form a unique fingerprint.

Cite this