Back to Search Document details
39th Meeting: Daejeon, KR, March 2025 2025-05-30 10:15
EE1-related: Deep Reference Frame Generation for Inter Prediction Enhancement
Abstract
This contribution reports the EE1-related test results of NN-based inter prediction. This contribution primarily focuses on an updated strategy for incorporating deep reference frames into the reference picture lists, while adopting the network architecture, training, and inference strategies from JVET-AF0208 [1]. Unlike JVET-AF0208, which considers the generated frame as an additional reference directly inserted into the reference picture list, the updated strategy replaces one of the existing reference frames with the generated reference frame. This ensures that the number of reference frames for B-frames and P-frames remains consistent with the default configuration of NNVC CTC.
JVET-AL0184 EE1-related: Deep Reference Frame Generation for Inter Prediction Enhancement [D. Ding, X. Chen, Z. Chen (Wuhan Univ.)]

This contribution reports the EE1-related test results of NN-based inter prediction. This contribution primarily focuses on an updated strategy for incorporating deep reference frames into the reference picture lists, while adopting the network architecture, training, and inference strategies from JVET-AF0208. Unlike JVET-AF0208, which considers the generated frame as an additional reference directly inserted into the reference picture list, the updated strategy replaces one of the existing reference frames with the generated reference frame. This ensures that the number of reference frames for B-frames and P-frames remains consistent with the default configuration of NNVC CTC.

Implemented on the VTM-11.0_NNVC-12.0, the test results under the NNVC-RA configuration (with LOP filter and NN-intra enabled) are reported as follows:

DRF network from JVET-AF0208 (3800K, 504kMAC/pixel): -2.38%/-1.28%/-1.41% bitrate savings for the Y/U/V components in RA.

The last frame is replaced, and the new frame goes to position 2 of the list.

Vimeo-90K triplet is used for training

Compared to JVET-AF0208, some loss occurs

Complexity approx. 500 kMAC/pix

SADL, no integer

It was commented that better results might be achieved if longer sequences would be used (triplet has only 3 frames), and/or more sophisticated training strategy would be applied like in current EE1.

Applied only to temporal layers 3-5 with closest distances.

It was agreed to study this in in EE1-2 with comparable condition in training etc. as the other method from Xidian Univ. Varying the number of reference pictures is not of highest importance in that study, but must be comparable for different proposals.

Decisions
It was agreed to study this in in EE1-2 with comparable condition in training etc. as the other method from Xidian Univ. Varying the number of reference pictures is not of highest importance in that study, but must be comparable for different proposals.
Citation