JVET-T0058 AHG11: Information on inter-prediction coding tool with deep neural network [B. Choi, Z. Li, W. Wang, W. Jiang, X. Xu, S. Liu (Tencent)]
This contribution was discussed at 0840 on Thursday 8 October (chaired by GJS & JRO).
This informative contribution reports some preliminary results of deep neural network (DNN) utilization for inter-prediction. The idea is inserting of a new (virtual) reference picture, which is generated by inference processes of trained networks, into reference picture list (RPL). The entire networks consist of several sub-network models for edge-detection, optical flow estimation/compensation and detail enhancement. Since networks were trained with small size training sequences, classes C&D show coding gains of 2.18%, 3.93% respectively in luma for RA, in comparison to other classes.
This basically uses the concept of frame-rate up-conversion (FRUC). A virtual reference picture is generated by NN inference and is inserted into the RPL. The extra picture is treated as a LTRP in the DPB, with MVs set to 0.
Proposed neural network process for generation of virtual reference picture
BD-rate and Coding Time for NN-based Video Coding Tool Testing in Inference Stage (Tested 33frames with GOP=32 and patch size 512x512)
Random Access Over VTM-10.0
BD-rate (Y)
BD-rate (U)
BD-rate (V)
Encoding Time (CPU)
Decoding Time (CPU)
Encoding Time (GPU)
Decoding Time (GPU)
Class A1
Class A2
Class B
−0.68%
1.43%
1.52%
133.82%
45103.39%
Class C
−2.18%
−0.27%
1.37%
116.60%
34643.88%
Class E
O
O
O
O
O
O
O
Overall
Class D
−3.93%
4.27%
6.51%
128.06%
59024.62%
Class F
The largest gain was for class D. The scheme was trained on low resolution video. However, even for class D, the numbers for chroma seem to indicate loss.
The entire processing chain consists of edge detection, optical flow estimation/compensation, and detail enhancement by NN.
The model was reportedly trained using lower resolution video, and the proponent said it might do better for higher resolution if trained for that case. There could be a library of such separately tuned models.
The work was said to be somewhat preliminary, and provided for information.
It was commented that in VVC we have made tools operate, e.g., on a 4x4 level rather than a single-sample level, and some of the gain is probably just coming from allowing sample-by-sample processing rather than block-based processing.
It was suggested to consider the interaction with other coding tools such as BDOF and DMVR, and with higher-complexity variations of those other methods (which were exercised e.g. in JEM) as anchors.
Adding other anchor methods of reference picture generation was suggested.
The proponent said that this scheme does not seem to suffer from blurring that is found in other methods.
Further study was encouraged.