Abstract
Interpretability of black-box deep models is yet challenging because existing model-agnostic methods mainly locally explain the behavior of the classifier by learning a linear proxy around the instance being predicted. The explanation can be faithful locally, but may not be accurate globally. In this paper, we for the first time formulate the interpretation of classifiers as a bandit problem and introduce a Bandit Interpretation method via Confidence Selection (BICS). We statistically impose disturbances on different arms (image regions) and examine non-linear changes of the model's output to fairly select important regions via Upper Confidence Bounds (UCB). Unlike previous model-agnostic methods that directly occlude super-pixels, our method softly applies perturbations at a pixel level and thus can fully explore more regions with multiple granularities, leading to a more precise and robust interpretation. Quantitative and qualitative experimental results demonstrate that our approach provides reasonable and precise explanations for various image recognition tasks on different models.
| Original language | English |
|---|---|
| Article number | 126250 |
| Journal | Neurocomputing |
| Volume | 544 |
| DOIs | |
| State | Published - 1 Aug 2023 |
Keywords
- Deep neural networks
- Interpretability
- Multi-armed bandit problem
- Statistical perturbation
- The upper confidence bounds
Fingerprint
Dive into the research topics of 'Bandit Interpretability of Deep Models via Confidence Selection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver