JVET-AF0075 Evaluation results of Low Complexity Enhancement Video Codec (LCEVC) with HM and VTM on 4K content [O. Chubach, Y.-L. Hsiao, C.-Y. Chen, C.-W. Hsu, T.-D. Chuang, Y.-W. Chen, Y.-W. Huang, S.-M. Lei (MediaTek)]
JVET-AF0075 to JVET, m65560 to WG 4 (cross-checked in JVET-AF0224, JVET-AF0228, JVET-AF0277)
This contribution presents evaluation results of Low Complexity Enhancement Video Codec (LCEVC) using its reference implementation (LTM-5.4.1) when combined with HEVC (HM-16.25) and VVC (VTM-19.0) as base layer codecs. Tests were performed for 4K content (Class A1 and A2 test sequences), following JVET Random Access (RA) Common Testing Conditions (CTC). The anchor is the result of single layer coding, i.e., encoding 4K videos by VTM or HM only, and the test is the result of base layer codec, which encodes quarter-resolution version of 4K video by using VTM or HM as the base layer, followed by enhancement layer coded using LTM-5.4.1, where the enhancement layer is obtained as a difference between original 4K video and upsampled reconstructed frames from the base layer. PSNR BD-rate was used for objective evaluation.
On top of VTM-19.0, the simulation results of LCEVC, VTM-19.0 + LTM-5.4.1,were reported as:
- RA CTC: {Y BD-rate = +31.59%, Cb BD-rate = +99.55%, Cr BD-rate = +70.60%}
On top of HM-16.25, the simulation results of LCEVC, HM-16.25 + LTM-5.4.1, were reported as:
- RA CTC: {Y BD-rate = +7.89%, Cb BD-rate = +78.01%, Cr BD-rate = +51.42%}
Visual artefacts were also reported and illustrated in the contribution using still-frame captures from the sequences. Subjective viewing of the results reportedly showed some loss of details in some LCEVC encoded areas.
m65556 (partly in JVET-AF0292, which was to be revised with the additional material from m65556) Considerations on LCEVC performances in relation to input document JVET-AF0075/m65560 [Simone Ferrara, Lorenzo Ciccarelli, Guendalina Cobianchi, Guido Meardi]
This document comments on the methodology and the conclusions of input document JVET-AF0075 and the expected performance of MPEG-5 Low Complexity Enhancement Video Coding (LCEVC) on objective metrics vs. formal MOS. Prior documents (about 30) were highlighted including LCEVC’s verification tests (m54455) and several evaluations of LCEVC using the same sequences (such as WG 11 N 19160, m53796, m52997 and m53523). A summary of LCEVC verification tests (VTs) can be found in WG 11 N19571. A relatively low correlation of PSNR with subjective quality was noted. The selection of bit rates for testing was also highlighted. The CTC for LCEVC, including configurations of LTM and anchors, was focused around a different range of bit rates, typically below 15 Mbps for UHD sequences – with the HM anchors ranging from QP 26–27 to QP 39–40, depending on the sequence – see also WG 11 N 18988. For the verification tests the QP values chosen for the anchors ranged between QP 27 and QP 42. Some results were shown with comparison to the x265 encoder for HEVC and the MainConcept encoder for VVC. Various metrics are discussed. It was emphasized that LCEVC has low complexity as a significant part of its design intent.
It was emphasized that PSNR has a relatively low correlation with subjective quality; subjective quality is what really matters. This contribution discussed various other objective metrics as well as PSNR, and in some cases those other metrics tend to have better correlation with subjective quality.
Chroma edges are less impactful on visual quality but the chroma fidelity balance with luma quality can be changed.
It was commented that the bit rate allocation between the base layer and enhancement layer, including temporal effects, is very important and has often not been optimized when performing comparisons. This includes per-frame, per-layer, bit rate allocation issues.
It was asked what kind of quantization was performed for the LCEVC VTs. The default quantization matrices (QMs) were reported to have been used for LCEVC. These are not flat, but are fixed-value matrices. It was said that flat matrices may have been used for the HEVC and VVC references. It was emphasized that although LCEVC used a non-flat matrix, it was a fixed-value matrix defined as the default rather than something customized by the encoder.
It was asked if dithering filtering was used in the verification test, and reported that it was not used.
There was a comment about x265 in relation to the HEVC HM encoder, saying x265 had significantly lower compression performance. It was commented that the x265 encoder configuration settings that were used in the tests reported in this contribution results in a large loss compared to the HM (in fact more than 80% loss in luma, more than 100% loss in chroma) and is worse than the AVC JM, saying having gains on top of a weak encoder may not compensate for the loss to the results obtained with a stronger single-layer encoder.
Another participant responded that what should be analysed is the delta between the layered and non-layered coding using the same base encoder rather than the relative performance with different base encoders. However, it was remarked that the delta does not seem to be invariant to the question of which base layer encoder is used.
It was commented that enforcing the use of three colour planes in LTM5.4.1 results in worse results compared to the tests reported in JVET-AF0075. Using a newer LTM version has different results but still was reported to have a loss.
It was remarked that carefully considering and studying the quality of an upscaled base layer encoding is very important. At low bit rates, an upscaled lower-resolution encoding may actually look better than (or as good as) a dual-layer or single-layer encoding of higher resolution video. It was reported that such an upscaled reference had been included in the LCEVC verification tests.
It was commented that VMAF is a trained metric; only luma was used for that training, and as a result, VMAF only considers the luma distortion.
It was asked whether the settings used for the LCEVC LTM5.4.1 encoder were used according to the settings reported in the verification testing, and the contributors of JVET-AF0075 confirmed that the settings were the same as reported in the WG 4 verification test report on the compression performance of LCEVC (WG 4 N 0076). This is also confirmed for the other contribution.
It was indicated that in JVET-AF0075 the quantization parameter (QP) values configured in its test followed the formula provided in the verification test report.
It was reported that the bitstreams from JVET-AF0075 were available and could be viewed upon request.
It was indicated that there are two enhancement layers in LCEVC, one of which is an upscaling layer.
It was commented that some parts of the video have better visual quality in the LCEVC layered coding approach than in the single-layer encodings, pointing to examples provided in the m65556 input contribution.