Back to Search Document details
36th Meeting: Kemer, TR, November 2024 2024-10-31 20:09
[AHG11] Towards Incorporating End-To-End Learning-Based Image Coding into Traditional Codec
Abstract
In this contribution, we present a method to combine end-to-end trained image compression methods with long-proven methods from VTM. The method comprises using an end-to-end coded image as reference image for a modified I-Frame. The encoder can then decide whether and how to use the AI-reference frame. A major advantage of this method is that data transfer from a GPU/NPU is only necessary on full frame level, thus solving a major problem in realizing neural-network-based components in hardware while using specialized hardware. This contribution aims at providing a proof-of-concept for this method and shows one possible way, E2E learning-based coding can be incorporated into traditional coding. With this hybrid method, we obtain rate-savings of -7.2% in Class E and -4.97% in Class C. For class C encoding run time overhead is 5% on CPU and 0.05% if GPU is used for E2E learning-based picture coding. This complexity-performance trade off is making this a promising new direction to explore, while keeping the main functionality of hybrid coders.
JVET-AJ0208 [AHG11] Towards Incorporating End-To-End Learning-Based Image Coding into Traditional Codec [F. Brand, T. Solovyev, E. Alshina (Huawei)]

In this contribution, a method is presented to combine end-to-end trained image compression methods with long-proven methods from VTM. The method comprises using an end-to-end coded image as reference image for a modified I-Frame. The encoder can then decide whether and how to use the AI-reference frame. A major advantage of this method is that data transfer from a GPU/NPU is only necessary on full frame level, thus solving a major problem in realizing neural-network-based components in hardware while using specialized hardware. This contribution aims at providing a proof-of-concept for this method and shows one possible way, E2E learning-based coding can be incorporated into traditional coding. With this hybrid method, we obtain rate-savings of -7.2% in Class E and -4.97% in Class C. For class C encoding run time overhead is 5% on CPU and 0.05% if GPU is used for E2E learning-based picture coding. This complexity-performance trade off is making this a promising new direction to explore, while keeping the main functionality of hybrid coders.

Contribution for information – no specific action requested.

It was commented that this could either be realized by using a P frame as I frame enhancement or by inter-layer prediction with an external reference.

VTM used “as normal”, including loop filters.

It was commented that the complexity of the NN based reference adds up to the complexity of VTM, even when using GPU for NN decoding the run time will be higher.

Is there a mechanism to inhibit the NN related parts of the bitstream for cases where it is not used? How often would that happen? It was answered that this highly depends on QP, at low QP it is hardly used and can then be disabled at frame level. NN codec uses relatively large blocks which could be switched on/off, so the granularity is coarse.

Results were only for AI, but in principle this could also be used for I pictures in RA. Gain less than 1% in that case (only class C and D reported).

Decisions
Results were only for AI, but in principle this could also be used for I pictures in RA. Gain less than 1% in that case (only class C and D reported).
Citation