JVET-J0021 Description of SDR, HDR and 360° video coding technology proposal by Qualcomm and Technicolor – low and high complexity versions [Y.-W. Chen, W.-J. Chien, H.-C. Chuang, M. Coban, J. Dong, H. E. Egilmez, N. Hu, M. Karczewicz, A. Ramasubramonian, D. Rusanovskyy, A. Said, V. Seregin, G. Van Der Auwera, K. Zhang, L. Zhang (Qualcomm), P. Bordes, Y. Chen, C. Chevance, E. François, F. Galpin, M. Kerdranvat, F. Hiron, P. de Lagrange, F. Le Léannec, K. Naser, T. Poirier, F. Racapé, G. Rath, A. Robert, F. Urban, T. Viellard (Technicolor)]
This contribution was discussed Wednesday 12 April 1910–1940 (chaired by GJS and JRO).
The non-360°, non-HDR aspects were presented.
This contribution describes the Qualcomm and Technicolor joint proposal in response to the CfP. The proposal contains majority of the tools that have been adopted into the JEM software. Additional or modified aspects include:
Triple-tree (TT) and asymmetric binary-tree (ABT) partition types (cf. JVET-D0117, JVET-D0064)
Various modifications of intra prediction and its mode coding (cf. JVET-D0113, JVET-D0119, JVET-D0110, JVET-H0057, JVET-D0114, JVET-G0060)
Merge, AMVP and affine motion are modified
Motion compensated padding
More transform choices, restriction of NSST usage (cf. JVET-C0022, JVET-D0126)
Sign prediction (cf. JVET-D0031)
Modified CABAC probability estimation (cf. JVET-G0112, JVET-E0119)
Filtering modifications
For HDR, pre-/post-dynamic range adaptation is used. For 360° video, ACP with geometric padding is used as a coding tool.
Objective SDR gains of 43.1% and 15.5% in terms of average luma BD-rate improvement have been reportedly achieved for constraint set 1 (i.e., RA) in high complexity mode, relative to HM and JEM anchors, respectively. For constraint set 2 (i.e., LD), the average luma BD-rate improvements are reportedly 33.7% relative to the HM anchor and 12.7 % relative to the JEM anchor. For this configuration, the encoder is about 1.5× as slow as the JEM and the decoder is about 16% faster.
In the low complexity mode, SDR gains of 39.7% and 10.3% in terms of average luma BD-rate improvement have reportedly been achieved for constraint set 1 relative to HM and JEM anchors, respectively. For constraint set 2, the average luma BD-rate improvements are reportedly 31.7% relative to the HM anchor and 9.9 % relative to the JEM anchor in low complexity mode. For this configuration, the encoder is about 2× the speed of the JEM and the decoder is about 15% faster than the JEM.
In the presentation, some other possible configurations were considered, e.g., modifying only the tree structure or disabling some features.
The software memory usage was about half that of the JEM, and lower than for the HM.
The software was a redesigned JEM, with substantial cleanup and an ability to disable individual tools relative to basically an HM core.
The low complexity configuration is without TT and ABT, using plain QTBT.
The software is a re-design of the JEM, with significantly reduced encoder (and decoder) run time.
HDR aspects were presented Friday 13 April 1225–1255 (chaired by JRO).
The additional document JVET-J0067 relates to HDR aspects of the proposal. From abstract of JVET-J0067:
This contribution provides additional information on the HDR video coding technology proposals by Qualcomm and Technicolor presented in JVET-J0021 and JVET-J0022. The proposed HDR technology is a colour volume transform (CVT) which is applied in the Y′CbCr 4:2:0 sample domain. The CVT is implemented through a Dynamic Range Adjustment (DRA) process which is applied as pre-processing at the encoder side, with the aim of improving the coding efficiency. At the decoder side, the inverse DRA process is applied.
Simulation results in this document reportedly show that the proposed CVT implemented on top of JEM7.0 software and tested on Class HDR-B test sequences provides around 34.0% and 8.3% of bit rate reduction (for PSNR-L100 metrics) against HM and JEM HDR anchors of the CfP, respectively. As it is shown in JVET-J0021, the proposed CVT being integrated in the core technology of JVET-J0021 (high complexity mode) provides for class HDR-B on average 41.3% and 18.8% BD-rate gain (PSNR-L100) against HM and JEM HDR anchors, respectively. In the low complexity Mode, the proposed CVT provides for class HDR-B on average 38.8% and 15.2% BD-rate gain (PSNR-L100).
Additionally, this document reports, that for for HDR-B class sequences proposed CVT utilized in the JVET-J0021 core design provides 14% of bit rate reduction (for PSNR-L100 metric) over the default (SDR) coding configuration in the high complexity mode and 13.6% of bit rate reduction over the SDR configuration in the low complexity mode.
HDR specific aspects:
A colour volume transform (CVT) including dynamic range adaptation (DRA) and cross-component DRA (outside of coding loop)
A lookup table that includes consideration of optimized chroma QP offset.
PSNRL100 and DeltaE100 were used for optimization of the proposal, and show similar objective gain over the JEM and HM anchors, which were optimized for wPSNR. In terms of wPSNR, the luma gain seems larger, but significant loss was observed in chroma. It was commented that it might be useful to compare this with the subjective results.
The HDR-related aspects of the proposal could be implemented outside the coding loop (e.g. controlled by an SEI message). For the submission, a fixed CVT is used over all sequences of a HDR category (PQ/HLG), but it could be also made sequence adaptive.
Comments from the discussion included:
It was noted that the balance between luma and chroma is shifted, relative to the JEM, with more improvement of luma than chroma.
In the two primary configurations that were presented, the decoder speed was about the same; the main change is in the encoder complexity. Another participant commented that there were some differences in complexity other than speed.
Lower complexity modes were also shown, illustrating a broader range of encoding and decoding compression-versus-speed tradeoffs.
360° related aspects were presented Friday 13 April 1715–1725 (chaired by JRO).
Dedicated tools for 360° video included:
Adjusted Cubemap Projection (ACP) is used.
Padding is added to the reconstructed cube faces and is symmetric around each cube face with width 64 samples.
The padded samples are obtained based on the ACP geometry and nearest-neighbour rounding.
The reconstructed ACP pictures are padded one time prior to in-loop filtering.
The padded reconstructed ACP pictures are sequentially processed by the deblocking filter, SAO and ALF before storage as reference pictures.
Motion compensated padding and OBMC for blocks on the boundary between top and bottom row of cube faces are disabled.
The padding area is 64 samples.
It is reported that the gain (on average) of 360°-specific tools is 2.3%, mainly due to padding. Padding is performed in the reference frame.