JVET-AO0173 [AHG11] A Hybrid Framework Integrating End-to-End Learned Intra-Frame Codec with Conventional Codec [N. Zou, A. B. Koyuncu, A. Hallapuro, F. Cricri, H. Zhang, J. Ahonen, M. M. Hannuksela (Nokia)]
This contribution proposes a hybrid framework that integrates end-to-end learned intra-frame compression (NLIC) methods with conventional compression techniques. The framework involves using NLIC-coded intra frames and VTM-coded inter frames. Furthermore, for each intra frame, the encoder decides whether to code it with NLIC or with VTM. Two sets of results are provided, depending on whether the NLIC was optimized by means of perceptual finetuning (PFT). With this hybrid framework, under the Random-Access configuration, the resulting BD-rates over NNVC-15.0 VTM (with NN tools off) anchor are reported to be as follows:
The system optimized with MSE and rate losses (no perceptual fine-tuning):
Overall -0.66% (Y), -2.90% (Cb), -1.31% (Cr)
Additionally, the Random-Access simulation results are evaluated with 8 perceptual metrics, the resulting perceptual BD-rates over NNVC-15.0 VTM (with NN tools off) anchor are reported to be as follows:
AVG | msssim Torch | Vif | Fsim | nlpd | iw-ssim | vmaf | psnrHVS | lpips | |
W/O PFT | -1.57% | -1.87% | -2.17% | -1.15% | -0.90% | -1.57% | -0.64% | -0.66% | -2.69% |
W/ PFT | -5.59% | -8.48% | -4.57% | -3.33% | -3.18% | -6.91% | -1.16% | -1.24% | -11.49% |
It was noted that there is no gain in AI – proponents explain that the model was not optimized for that case (models were optimized for different QP), such that the encoder always decides for VVC intra.
The algorithm is the same as in EE1-6.
It was asked what the benefit would be compared to the multi-layer framework (saving the HLS overhead)?
This cannot be directly compared, because the switching mechanism has not been used in the EE.
One option would be to investigate in the multi-layer framework what the benefit of the switching mechanism at picture level would be (which requires duplicate encoding AI/VVC), and analyze from the bitstream how large the HLS overhead is.
It was commented that this could later be extended to local switching mechanisms.
It was concluded to investigate this in EE1 in the context of the multi-layer interface.