Back to Search Document details
23rd Meeting: by teleconference, July 2021 2021-07-07 03:59
CE-related: On constraining inverse transform precision for high bit depth and high bitrate video coding
Abstract
At the 22nd JVET meeting extended precision mode was adopted to VVCv2 draft specification for high bit depth and high bitrate video coding. In this mode, the precision of the inverse transform is extended to 19 bits for coding of 12 bits video and to 21 bits for coding of 16 bits video. This contribution proposes to employ Content Adaptive Transform Precision (CATP) method to constrain the bit depth of inverse transform implementation while preserving coding performance of the Anchor (VTM13.0, extended precision ON). The method proposed here is a modification of the CATP of JVET-W0050 (CE2.2) and targets complexity reduction and constraining precision used in the inverse transform.
JVET-W0092 CE-related: On constraining inverse transform precision for high bit depth and high bit rate video coding [L. Kerofsky, D. Rusanovskyy, M. Karczewicz (Qualcomm), K. Naser, F. Galpin, F. Le Leannec, P. de Lagrange (Interdigital)]

This contribution was presented July 7 at 2330-0010.

At the 22nd JVET meeting extended precision mode was adopted to VVCv2 draft specification for high bit depth and high bit rate video coding. In this mode, the precision of the inverse transform is extended to 19 bits for coding of 12 bits video and to 21 bits for coding of 16 bits video. This contribution proposes to employ Content Adaptive Transform Precision (CATP) method to constrain the bit depth of inverse transform implementation while preserving coding performance of the Anchor (VTM13.0, extended precision ON). The method proposed here is a modification of the CATP of JVET-W0050 (CE2.2) and targets complexity reduction and constraining precision used in the inverse transform.

Method has been tested under HBD / HBR CTC with 16-bit, 17-bit and 18-bit constraints vs. the Anchor. No run time increase was reported for tested configurations, no noticeable penalty is reported for Normal QP configuration. Summary of the results for Low QP are following:

Test on 16 bit constraints:

Low QP

HDR PQ12

HDR HLG12

SVT12

SVT16

wY

wU

wV

Y

U

V

Ave. GBR

Ave. GBR

AI

-0.01%

-0.01%

-0.02%

0.00%

-0.01%

-0.01%

0.03%

0.28%

LDB

-0.02%

0.01%

0.00%

0.00%

-0.02%

-0.02%

0.01%

0.14%

RA

0.00%

-0.01%

0.00%

0.00%

-0.01%

0.00%

0.01%

0.15%

Test on 17 bit constraints:

Low QP

HDR PQ12

HDR HLG12

SVT12

SVT16

wY

wU

wV

Y

U

V

Ave. GBR

Ave. GBR

AI

-0.01%

-0.01%

-0.03%

-0.01%

-0.01%

-0.01%

0.01%

0.10%

LDB

0.00%

-0.01%

0.01%

0.00%

-0.01%

-0.01%

0.00%

0.03%

RA

0.00%

-0.01%

0.00%

0.00%

0.00%

-0.01%

0.00%

0.04%

Test on 18 bit constraints:

Low QP

HDR PQ12

HDR HLG12

SVT12

SVT16

wY

wU

wV

Y

U

V

Ave. GBR

Ave. GBR

AI

-0.01%

-0.03%

-0.03%

0.00%

-0.01%

-0.01%

0.00%

0.02%

LDB

-0.01%

-0.01%

0.01%

-0.01%

-0.02%

0.00%

0.00%

-0.01%

RA

0.00%

-0.01%

0.01%

-0.01%

-0.01%

0.00%

0.00%

0.00%

It was proposed to adopt proposed CATP method and constrained bit depth implementation of the inverse transform to VVCv2.

The proponents report that the extended precision flag of current v2 draft increases buffer size of transform coefficients by 18% (for 12 bit data), whereas the proposal does not have any increase in buffer size.

For 16 bit data, the increase of buffer size by the extended precision flag is larger (about 30%, from 16 to 21 bits). When keeping 16 bit internal bit depth, the proposal has a compression loss of roughly 0.3% for AI, 0.15% for RA/LB. When the proposal is using 18 bit internal bit depth, there is no loss for 16 bit data, but the buffer size is increased by 12.5% relative to VTM; but decreased by 14% compared to extended precision flag.

In order to achieve that, it is necessary to parse the entire TU and determine the maximum transform coefficient in that TU before the final inverse quantization and inverse transform can be done.

The absolute saving in memory (considering the cases where no loss occurs) is 3 bits per coefficient, which gives 768 bytes for the entire 64x32 block (considering the zero-out of secondary transform).

Implementation-wise, it would be either necessary to add another buffer for the inverse quantization (which would end up with more buffer requirement than for the usage of extended precision flag), or perform inverse quantization twice (in which latter case, the latency would be increased).

No action was taken on this.

Decisions
No action was taken on this.
Citation