Back to Search Document details
21st Meeting: by teleconference, January 2021 2021-01-07 09:36
A DNN Architecture for Intra-Frame Coding in YUV 4:2:0 format with Cross-Component Prediction
Abstract
Most existing deep neural network (DNN) based video coding architectures are designed to operate in non-subsampled input formats such as RGB or YUV 4:4:4. However, video coding standards such as HEVC and VVC are designed primarily for the YUV 4:2:0 colour format. In the 20th JVET meeting, JVET-T0123 introduced three distinct network architectures including a solution jointly coding luma and chroma components. This contribution document proposes two alternative joint coding designs for intra-frame coding in YUV 4:2:0 colour format. The experimental results demonstrate that the proposed methods outperform HM-16.20 about 12.5% in luma coding. Over the joint coding scheme in JVET-T0123, Method 1 provides 6.43% Y, 5.56% U, 6.76% V, and Method 2 achieves 6.64% Y, 6.97% U, 8.34% V coding improvements.
JVET-U0079 A DNN Architecture for Intra-Frame Coding in YUV 4:2:0 format with Cross-Component Prediction [H. E. Egilmez, A. K. Singh, M. Coban, M. Karczewicz (Qualcomm)]

Most existing deep neural network (DNN) based video coding architectures are designed to operate in non-subsampled input formats such as RGB or YUV 4:4:4. However, video coding standards such as HEVC and VVC are designed primarily for the YUV 4:2:0 colour format. At the 20th JVET meeting, JVET-T0123 discussed three distinct network architectures including an approach jointly coding the luma and chroma components. This contribution document proposes two alternative joint coding designs for intra-frame coding in the YUV 4:2:0 colour format. The experimental results reportedly demonstrate that the proposed methods outperforms HM-16.20 by about 12.5% in luma coding. Over the joint coding scheme in JVET-T0123, Method 1 reportedly provides 6.43% Y, 5.56% U, 6.76% V coding improvements, and Method 2 reportedly provides 6.64% Y, 6.97% U, 8.34% V coding improvements.

The Y/U/V performance compared to HM-16.20 is reported to be −12.41%/46.07%/13.65% and −12.61%/37.48%/10.91%, i.e., luma has coding gain but chroma has coding loss.

Compared to JVET-T0123, this contribution has higher coding gain in all 3 colour components, with a somewhat reduced number of parameters.

It was noted that the ratio of luma weight and chroma weight in the loss function is 4x.

Regarding the chroma coding performance loss, it was asked if the loss is reflected in visual quality. It was reported that the average loss is mainly due to only a few test sequences, particularly ParkingRunning and Campfire. And the proponent reported that the visual quality of these sequences appeared to be OK.

Regarding the observation that the GDN can be removed without performance loss (method 2 vs. method 1), it was commented that the additional 1x1 convolution layer in this contribution on top of JVET-T0123 seems to make GDN unnecessary.

It was commented that bit allocation between different colour components may be needed to address the large chroma loss for some test sequences.

Further study of YUV 4:2:0 coding in the context of the E2E framework is encouraged.

Decisions
Further study of YUV 4:2:0 coding in the context of the E2E framework is encouraged.
Citation