JVET-AE0191 AhG11 - EE1-0 High Operation Point model [F. Galpin (InterDigital), S. Eadie, D. Rusanovskyy (Qualcomm), Y. Li, J. Li (ByteDance), L. Wang, R. Chang (Tencent), Z. Xie (Oppo), E. Alshina (Huawei)] [late]
A Unified Filter Architecture for the High-performance Operation Point (HOP) was defined at JVET-AD0380 and includes filter architecture, training process and training test set. Every design element and training procedure for this architecture were carefully selected though multiple rounds of earlier EE1 tests. Following JVET-AD meeting, a unified dataset and unified training procedure was implemented by multiple participants, and training conducted in parallel independently to allow cross-check.
This was initially discussed in JVET on July 13.
This contribution reports on the progress of unified filter implementation and achieved simulation results.
Intermediate results demonstrate the performance of the model: -7.78%, -18.81%, -19.98% BD-rate change in Y, U and V components for AI configuration and -9.49%, -22.48%, -22.17% BD-rate change in Y, U and V components for RA configuration.
It is suggested to adopt EE1-0 into NNVC common software and enable it by default for HOP. It is also suggested to make training strategy defined for EE1-0 a part of NNVC SW. It is also suggested to continue practice of joint training activity for most promising NN-based tools studied in EE1. Identified dis-convergence of training at Stage II can be one of aspects to study in further EE1 tests.
In stage 2, problem with switch from L1 to L2 loss function: In the first epoch (18), the loss is still decreasing, and then it increases. It was suggested that a solution might be reduction of the learning rate. Results above are with the model coming from epoch 18 of stage 2.
In each stage, training is performed from scratch. Currently, training of stage 3 is running, which uses the filter model from stage 2 only for generating decoded data, but then a completely new model is trained from scratch.
It was agreed that a unified training strategy shall be used in EE to make proposals in a given category comparable. In the same category (e.g., directly competing methods, such as filter for HOP), identical number of epochs shall be used, training parameters (loss function, batch size, switching point, learning rate decay) shall be identical.
For proposals that are exclusive (e.g., HOP and LOP loop filters), the following shall apply:
- It would be desirable (to be clarified in BoG if mandatory already in the next round of EE) to define an equivalent mandatory strategy for LOP (the current LOP was trained somewhat similarly). Data set should be identical, but number of epochs may be different, as well as training parameters.
- not only the data set, but also augmentation, generation of coded data (QP points) should be identical.
Proponents are free to provide additional results with a different training strategy.
In terms of the HOP model for next NNVC SW the following was agreed.
Decision:
- Inference results for integerized stage 2 are expected during the meeting, such that a decision could be made. D. Rusanovskyy will send the model to cross-checkers, and provide an input when results are available. This could become the “fallback candidate” for HOP in NNVC6 SW.
- Training for stage 3 is running (plus cross-checks), but will not be available before the end of the meeting. An integerization strategy for stage 3 should be negotiated in the BoG. If the integerized stage 3 model provides better results within 14 days after the meeting, it will replace the “fallback candidate” in NNVC6 (to be confirmed in telco on Aug. 2)
It was later agreed (JVET Monday 17 July afternoon session, see notes under JVET-AE0291) to shift the deadline above and Telco by one week, August 9.