JVET-AL0243 [AHG11] Multilayer framework for supporting a hybrid codec using End-to-End Learned Image Codec and Conventional Video Codec [F. Urban, Y. Chen, F. Galpin, E. François (InterDigital)] [late]
In this contribution, we present a framework to support a hybrid codec that combines End-to-End Learned (E2E) image compression methods with conventional compression methods such as VVC. The framework is based on a modification of VVC to allow external picture coding. The multi-layer capability of VVC can then be used transparently to allow the use of such pictures inside VVC, while allowing the encoder to do low-level RD choices in the enhancement layer. The hybrid codec involves using E2E-coded intra frames and VTM-coded inter frames. The main advantage of this approach is that it allows the use of an external picture without defining explicitly the external codec, while allowing the use of a standard VVC decoder on the enhancement layer. In practice, the external codec can be run independently on AI accelerator and only the resulting picture is used by the VVC decoder.
An example of use with JPEG-AI encoded intra pictures is demonstrated using random-access configuration and multi-layer on top of VTM 23.
A mechanism is employed similar to SHVC (such a mechanism is not existing in VVC) to use an external reference as base layer, integrated in NNVC software. JPEG-AI is used, bitstream is separate.
Unlike the other two proposals, the inter picture (if in layer 1 or higher) might itself select if it is better to use the AI coded (layer 0) or the conventional coded (layer 1) reference.
Further study in AHG11/14, with the goal to combine the aspects of the three methods of JVET-AL0196, JVET-AL0203 and JVET-AL0243 to provide an interface with NN-coded intra frames. As a first effort, proponents should discuss the options, benefits and priorities of elements of such an interfaces.