Skip to main navigation Skip to search Skip to main content

Adult image and video recognition by a deep multicontext network and fine-to-coarse strategy

  • Xinyu Ou
  • , Hefei Ling
  • , Han Yu
  • , Ping Li
  • , Fuhao Zou
  • , Si Liu
  • Huazhong University of Science and Technology
  • CAS - Institute of Information Engineering
  • Yunnan Open University

Research output: Contribution to journalArticlepeer-review

Abstract

Adult image and video recognition is an important and challenging problem in the real world. Low-level feature cues do not produce good enough information, especially when the dataset is very large and has various data distributions. This issue raises a serious problem for conventional approaches. In this article, we tackle this problem by proposing a deep multicontext network with fine-to-coarse strategy for adult image and video recognition. We employ a deep convolution networks to model fusion features of sensitive objects in images. Global contexts and local contexts are both taken into consideration and are jointly modeled in a unified multicontext deep learning framework. To make the model more discriminative for diverse target objects, we investigate a novel hierarchical method, and a task-specific fine-to-coarse strategy is designed to make the multicontext modeling more suitable for adult object recognition. Furthermore, some recently proposed deep models are investigated. Our approach is extensively evaluated on four different datasets. One dataset is used for ablation experiments, whereas others are used for generalization experiments. Results show significant and consistent improvements over the state-of-the-art methods.

Original languageEnglish
Article number68
JournalACM Transactions on Intelligent Systems and Technology
Volume8
Issue number5
DOIs
StatePublished - Jul 2017
Externally publishedYes

Keywords

  • Adult image and video recognition
  • Deep convolutional network
  • Fine-to-coarse strategy
  • Multicontext modeling

Fingerprint

Dive into the research topics of 'Adult image and video recognition by a deep multicontext network and fine-to-coarse strategy'. Together they form a unique fingerprint.

Cite this