Back to Search Document details
23rd Meeting: by teleconference, July 2021 2021-07-13 18:38
AHG11: Deep In-Loop Filter with Adaptive Model Selection and External Attention
Abstract
This contribution presents a convolutional neural network-based in-loop filtering method wherein adaptive model selection is introduced. The proposed Deep in-loop filter with Adaptive Model selection (DAM) method is developed from the prior contribution JVET-V0100, introducing a new network structure to the code base of VTM-11.0+NewMCTF. Compared with VTM-11.0+NewMCTF, the proposed method reportedly shows on average {9.12%, 22.39%, 22.60%}, {12.32%, 27.48%, and 27.22%}, and {10.61%, 28.19%, and 27.91%} BD-rate reductions for {Y, Cb, Cr}, under AI, RA, and LDB configurations, respectively.
JVET-W0100 AHG11: Deep In-Loop Filter with Adaptive Model Selection and External Attention [Y. Li, K. Zhang, L. Zhang (Bytedance)]

This contribution presents a convolutional neural network-based in-loop filtering method wherein adaptive model selection is introduced. The proposed Deep in-loop filter with Adaptive Model selection (DAM) method is developed from the prior contribution JVET-V0100, introducing a new network structure to the code base of VTM-11.0+NewMCTF. Similar to JVET-V0100, residual blocks are utilized as the basic module and stacked several times to construct the final network. As a further improvement from JVET-V0100, external attention mechanism is introduced in this contribution, leading to an increased representation capability with a similar model size. In addition, to deal with different types of content, individual networks are trained for different types of slices and quality levels. Compared with VTM-11.0+NewMCTF, the proposed method reportedly shows on average {9.12%, 22.39%, 22.60%}, {12.32%, 27.48%, and 27.22%}, and {10.61%, 28.19%, and 27.91%} BD-rate reductions for {Y, Cb, Cr}, under AI, RA, and LDB configurations, respectively.

Model selection is at slice or CTU level.

SAO and deblocking are disabled/replaced.

The attention map is computed from prediction and residual block data.

1.43 MMACs/pixel (more complex than the highest in EE, but also better performance); CPU decoder run time increased by approximately 1000x.

A reduced complexity version was also presented.

It was pointed out that the above analysis seems to be only for luma, chroma should be included.

It was pointed out that the models are switched at finer granularity (32x32), which also has complexity impact in terms of reloading model parameters.

24 models were used in total, but at 32x32 granularity only 3 can be selected from a candidate list (based on QP).

In training, the same augmentations of original blocks were used as in previous proposal JVET-V0100 (rotation, flip).

Decisions
In training, the same augmentations of original blocks were used as in previous proposal JVET-V0100 (rotation, flip).
Citation