Back to Search Document details
39th Meeting: Daejeon, KR, March 2025 2025-03-25 14:01
AHG11: Backbone Block Enhancement of LOP In-Loop Filter with Over-Parameterized Training and Variable Channels
Abstract
This contribution proposes backbone block (BBB) enhancement of LOP in-loop filter with over-parameterized training and variable channels. The 1x3 and 3x1 separable convolutions in the LOP BBB are replaced with the 1x5 and 5x1 separable convolutions to increase the receptive field and enhance the prediction accuracy of the LOP network. Over-parameterized training to extend the 1x5 and 5x1 separable convolutions into a multi-branch structure is utilized to enhance the multi-scale feature extraction capability of the LOP BBB. Variable channels are adopted in the LOP BBB to enhance the richness of its input information while reducing network complexity. Therefore, this contribution reduces the complexity of LOP4 in-loop filter from 16.83 kMAC/pixel to 16.81 kMAC/pixel while enhancing the BD-rate performance. Compared to the NNVC-11.0 anchor, the BD-rate performance of the proposed float model with NNIntra enabled provides: {-0.19% (Y), -1.57% (U), -1.70% (V)} for the AI configuration, and {-0.07% (Y), -2.42% (U), -2.44% (V)} for the RA configuration.
JVET-AL0137 AHG11: Backbone Block Enhancement of LOP In-Loop Filter with Over-Parameterized Training and Variable Channels [J. Han, C. Jung, Q. Qin (Xidian Univ.)]

This contribution proposes backbone block (BBB) enhancement of LOP in-loop filter with over-parameterized training and variable channels. The 1x3 and 3x1 separable convolutions in the LOP BBB are replaced with the 1x5 and 5x1 separable convolutions to increase the receptive field and enhance the prediction accuracy of the LOP network. Over-parameterized training to extend the 1x5 and 5x1 separable convolutions into a multi-branch structure is utilized to enhance the multi-scale feature extraction capability of the LOP BBB. Variable channels are adopted in the LOP BBB to enhance the richness of its input information while reducing network complexity. Therefore, this contribution reduces the complexity of LOP4 in-loop filter from 16.83 kMAC/pixel to 16.81 kMAC/pixel while enhancing the BD-rate performance. Compared to the NNVC-11.0 anchor, the BD-rate performance of the proposed float model with NNIntra enabled provides: {-0.19% (Y), -1.57% (U), -1.70% (V)} for the AI configuration, and {-0.07% (Y), -2.42% (U), -2.44% (V)} for the RA configuration.

It was commented that the proposal has interesting gains, to be studied in an EE. Similar comments apply as for JVET-AL0136 regarding aligning model, training, comparison point, etc.

Additional complexity impact of the increased filter lengths should be studied, as well as the contribution in gain that comes by the usage of 5x1/1x5 filter kernels.

It was asked how the kMAC number decreases while the number of parameters increases. This is caused by reducing BBB blocks to 32x32 earlier, while the longer filter kernels require more parameters.

Decisions
It was asked how the kMAC number decreases while the number of parameters increases. This is caused by reducing BBB blocks to 32x32 earlier, while the longer filter kernels require more parameters.
Citation