Back to Search Document details
30th Meeting: Antalya, TR, April 2023 2023-04-22 15:34
AhG11 : neural network-based intra prediction with reduced complexity
Abstract
Following the 29th JVET meeting, 11-20 January 2023, three additional neural network-based coding tools [1, 2, 3] have been integrated into VTM-11-NNVC [4]. In [1], a content-adaptive post-filter takes reconstructed luma and chroma samples of a given reconstructed frame, along with the slice QP, to return a filtered version of these samples. This post-filter consists of four stacks of convolutional layers trained offline. At the encoder side, the multipliers of one of these stacks, which are the coefficients multiplying the result of a convolution before applying the non-linearity, are fine-tuned on the current video sequence. Then, the multipliers update, i.e. the difference between the initial multipliers and their fine-tuned version, are coded using Neural Network compression and Representation (NNR) and sent to the decoder. In [2], two neural networks are used as upsamplers in RPR. For a given luma channel, its low-resolution reconstruction, its low-resolution prediction, the slice QP, and the base QP are fed into the first neural network, and the neural network output is added to the result of the RPR upsampling of the low-resolution reconstruction to obtain the high-resolution reconstruction. For a given pair of chroma channels, the neural network-based upsampling is almost identical to that in luma, but using the second neural network. In [3], a neural network-based intra prediction mode is added to the set of intra prediction modes in VVC.
JVET-AD0212 AhG11: neural network-based intra prediction with reduced complexity [T. Dumas, F. Galpin, P. Bordes (InterDigital)] [late]

Initial version rejected as “placeholder”.

Following the 29th JVET meeting, 11-20 January 2023, three additional neural network-based coding tools have been integrated into VTM-11-NNVC. Firstly, a content-adaptive post-filter takes reconstructed luma and chroma samples of a given reconstructed frame, along with the slice QP, to return a filtered version of these samples. This post-filter consists of four stacks of convolutional layers trained offline. At the encoder side, the multipliers of one of these stacks, which are the coefficients multiplying the result of a convolution before applying the non-linearity, are fine-tuned on the current video sequence. Then, the multipliers update, i.e. the difference between the initial multipliers and their fine-tuned version, are coded using Neural Network compression and Representation (NNR) and sent to the decoder. Second, two neural networks are used as upsamplers in RPR. For a given luma channel, its low-resolution reconstruction, its low-resolution prediction, the slice QP, and the base QP are fed into the first neural network, and the neural network output is added to the result of the RPR upsampling of the low-resolution reconstruction to obtain the high-resolution reconstruction. For a given pair of chroma channels, the neural network-based upsampling is almost identical to that in luma, but using the second neural network. Third, a neural network-based intra prediction mode is added to the set of intra prediction modes in VVC.

To limit the complexity of the neural network-based intra prediction mode, each neural network in this mode features a relatively small architecture and weight sparsity is enforced. However, the architecture and weight sparsity policy has not been tuned to each neural network yet. This contribution thus proposes a refined architecture and weight sparsity policy to reach a new tradeoff between mean BD-rate gains and complexity.

It is reported that, on top of VTM-11-NNVC-4, whereas the original neural network-based intra prediction mode yields -3.61%, -3.16%, -3.27% and -1.81%, -0.49%, -0.90% in AI and RA respectively for worst-case complexity of 7.8 kMACs/pixel, the neural network-based intra prediction mode with the refined policy yields -3.38%, -3.00%, -3.18% and -1.72%, -0.61%, -0.99% in AI and RA respectively for worst-case complexity of 4.8 kMACs/pixel.

Parameter memory is reduced from 5.6 MB to 4.9 MB.

It was asked if 4x4 is needed, as this still has highest number of kMAC/pixel.

It was agreed to investigate this in an EE, including a training crosscheck.

Decisions
It was agreed to investigate this in an EE, including a training crosscheck.
Citation