JVET-J0028 Description of SDR and HDR video coding technology proposal by Sony [T. Suzuki, M. Ikeda, K. Sharman (Sony)]
This contribution was discussed Thursday 12 April 1430–1505 (chaired by GJS and JRO).
This contribution presents a description of SDR and HDR video coding technology proposal by Sony in response to the CfP. The proposed techniques were developed on top of the JEM and the codec design is common between SDR and HDR. The proposed techniques (relative to JEM) are
Sign prediction
Use of multiple reference samples in intra prediction
Modified PDPC planar (part of which is a bug fix)
Transform matrix replacement (reducing the number of transforms from 5 to 2, but with flipping and transposing)
Adaptive multiple core transforms (for luma and chroma, with a flag to indicate whether the chroma is the same as for luma or is DCT2 variant)
Adaptive scaling for transform and quantization
Affine MC with reduced overhead, adaptively using a 3-parameter or 4-parameter model, with lowest block size either 4x4 (with 2x2 chroma) or 2x2 (with 1x1 chroma) – cases corresponding to translation, zoom, rotation and general affine
Large CTU up to 256x256 (with CBF set to 0 when the largest size used; JEM anchor is 128x128 max)
Extended deblocking filter (for large blocks and also for chroma)
Modified adaptive loop filter classification
There is no use of pre-processing outside the coding loop and no specific optimizing of encoding parameters was done using non-automatic means (e.g. on a per-sequence basis) in both SDR and HDR. Quantization settings are kept static except for a one-time change of the settings to meet the target bit rate.
The contribution reports a coding gain for Y, U and V, on average, of 2.41%, 4.85% and 5.1%, and 2.25%, 6.74% and 7.34% over JEM at SDR constraint set 1 and 2 (i.e., RA and LD), respectively. For HDR, it reports a coding gain for Y, U and V, on average, of 2.35%, 5.14% and 7.73%, and 1.78%, 6.69% and 8.89% over JEM at HDR-A and HDR-B constraint set 1, respectively.
Encoder runtime was about 4× of JEM, decoder was about 1.3× JEM. Proponents believe that the large increase of encoder runtime is mainly due to RDO with larger CTUs.
Comments from the discussion:
It was asked how often the large CTUs seem to be used. The proponent did not know. The gain for this might be in the neighbourhood of 1%, but hardware implementers are not fond of it. It was commented that the primary implementation problem is the maximum transform size rather than the maximum CTU size. Most of the benefit in coding efficiency was said to come from the large CTU size, not the large transform.
No further detailed presentation was needed on HDR, as there are no specific tools. The results above relate to PSNRY.
In the table below are complete results with all metrics.
Over HM
DE100
PSNRL100
wPsnrY
wPsnrU
wPsnrV
psnrY
psnrU
psnrV
Average HDR-A
−57.20%
−32.60%
−29.91%
−66.72%
−69.31%
−29.76%
−63.83%
−68.77%
Average HDR-B
−36.18%
−27.94%
−28.65%
−54.81%
−51.79%
−27.59%
−52.48%
−47.58%
Average all
−44.06%
−29.69%
−29.12%
−59.27%
−58.36%
−28.40%
−56.74%
−55.53%
Over JEM
DE100
PSNRL100
wPsnrY
wPsnrU
wPsnrV
psnrY
psnrU
psnrV
Average HDR-A
−6.51%
−2.88%
−2.44%
−5.86%
−7.84%
−2.35%
−5.14%
−7.73%
Average HDR-B
0.52%
0.34%
−0.73%
−4.18%
−6.02%
−1.78%
−6.69%
−8.89%
Average all
−2.12%
−0.87%
−1.37%
−4.81%
−6.70%
−1.99%
−6.11%
−8.46%