Back to Search Document details
37th Meeting: Geneva, CH, January 2025 2025-01-16 19:12
AHG 11: Neural Network Coded Reference Frame for Intra Coding
Abstract
This contribution presents a method to combine end-to-end trained image compression methods with long-proven methods from VTM. The method comprises using an end-to-end coded image as reference image for a modified I-Frame. The encoder can then decide whether and how to use the NN-reference frame. In this method data transfer from a GPU/NPU is only necessary on full frame level into the decoded picture buffer, thus solving a major problem in realizing neural-network-based components in hardware while using specialized hardware. Compared to the previous contribution JVET-AJ0208 the method was harmonized with LMCS, showing similar gains. This contribution presents two variants, one where the encoder checks for each frame, whether the NN-reference frame is beneficial and should be used and one where this decision is based purely on the QP. In the first case gains of -4.21%/-7.07%/-3.63% in AI can be achieved with currently 138% encoder runtime. In the second case the gains are -3.29%/-0.43%/0.10% with 49% encoder runtime.
JVET-AK0177 AHG 11: Neural Network Coded Reference Frame for Intra Coding [F. Brand, T. Solovyev, E. Alshina (Huawei)]

This contribution presents a method to combine end-to-end trained image compression methods with long-proven methods from the VTM. The method comprises using an end-to-end coded image as a reference image for a modified I-frame. The encoder can then decide whether and how to use the NN-reference frame. In this method data transfer from a GPU/NPU is only necessary on full frame level into the decoded picture buffer, thus solving a major problem in realizing neural-network-based components in hardware while using specialized hardware. Compared to the previous contribution JVET-AJ0208 the method was harmonized with LMCS, showing similar gains. This contribution presents two variants, one where the encoder checks for each frame, whether the NN-reference frame is beneficial and should be used and one where this decision is based purely on the QP. In the first case gains of -4.21%/-7.07%/-3.63% in AI can be achieved with currently 138% encoder runtime. In the second case the gains are -3.29%/-0.43%/0.10% with 49% encoder runtime.

Used for QP>22, otherwise VTM intra.

It was commented that more sequence adaptation might be better. In some cases, luma and chroma are quite unbalanced.

The “intra difference coding” was newly implemented in VTM. It was pointed out that multi-layer VVC could be used, however RDO would need to be modified and aligned between NN and VVC. Frame-level RDO is used, with first pass comparing cost of VVC and NN

Investigated with ECM? No.

It was asked whether RA results were available? Simulations were still ongoing.

It was askes why not JPEG-AI was used? That is not optimized for PSNR.

It was commented that it would be interesting also seeing results with other metrics.

Further study was requested.

Decisions
Further study was requested.
Citation