JVET-AK0146 [AHG11] A Hybrid Framework Integrating End-to-End Learned Image Codec with Conventional Codec [N. Zou, A. Hallapuro, F. Cricri, H. Zhang, M. M. Hannuksela (Nokia)]
This contribution proposes a hybrid framework that integrates end-to-end learned image compression (LIC) methods with conventional compression techniques. The framework involves using LIC-coded intra frames and VTM-coded inter frames. Furthermore, the encoder decides whether to use the LIC-coded intra frames or the VTM-coded intra frames. With this hybrid framework, under the Random-Access configuration, the resulting BD-rates over NNVC-7.1 VTM (with NN tools off) anchor were reported to be as follows:
Class A1 0,00% (Y), 0,00% (Cb), 0,00% (Cr)
Class A2 -0,27 % (Y), -0,34 % (Cb), -0,35 % (Cr)
Class B -0,18 % (Y), -8,81 % (Cb), -8,35 % (Cr)
Class C 0,22 % (Y), -7,60 % (Cb), -7,32 % (Cr)
Class D -0,70 % (Y), -11,48 % (Cb), -11,34 % (Cr)
This uses a very large network, with 3.3 MMAC/s encoder, 8.5 for decoder; the exact architecture was not described.
Training uses BVI, DIV2K, JPEG-AI, CLIC and some other
Proponents are requested to upload the presentation (in v2 another version was provided)
Encoder can switch between conventional and learned method for the entire I frame.
Further study was highly encouraged. A more precise description should be provided, and complexity reduction, with a reduction of number of models, was recommended.