JVET-V0090 AHG11: Neural network based temporal processing [B. Choi, Z. Li, W. Wang, W. Jiang, X. Xu, S. Liu (Tencent)]
This proposal reports the test results of a neural network (NN) based temporal filtering for both detail enhancement and inter-prediction. The proposed NN processing is composed of two steps, NN-based reconstruction process and NN-based prediction process. Both processes has the same NN model utilizing temporal features, but the parameters of each model was individually trained for improving the quality of the reconstruction or minimizing error of the motion compensated prediction. The test results show coding gains 3.14%, 4.69% and 6.87% respectively, for classes B, C & D in luma for RA.
No class A results yet.
Two stages of NN, each with temporal processing invoking additional reference pictures (using past and future pictures):
- Reconstruction/loop filter (before DPB)
- Prediction filter (after DPB, before MC)
The proponent asserts that this is mainly effective in case of small motion, but might have problems with large motion.
Would it make sense to use temporal processing only for the second stage, and employ another filter as loop filter? What are the gains of stage 1 and stage 2 individually? Has only roughly been investigated so far. Is there overlap in the gain of the two filters? Yes.
The current approach replaces one of the reference pictures used for prediction by the temporally processed version. It is pointed out that this might not be optimum.
The benefit of stage 2 is probably more interesting.
Further study was encouraged, but it would be premature to investigate in an EE.