JVET-AM0010 JVET AHG report: Encoding algorithm optimization (AHG10) [K. Andersson, P. de Lagrange, A. Duenas (co-chairs), T. Ikai, T. Solovyev, A. Tourapis (vice chairs)]
Related contributions
The following contributions were identified relating to AHG10 and are summarized in the following paragraphs.
JVET-AM0045: AHG17 Generated VTM anchor bitstreams using LambdaScaleTowardsNextQP
This document is a follow up the approach of scaling lambda to control target bitrate as suggested in JVET-AL0207 with the current approach for target bitrate matching where QP is increased by 1 at a selected frame and used for the remaining length of the video sequence. The document reports results of rate matching using VTM-23.9 where the QP closest to the target bitrate is selected and the lambda is scaled to increase or decrease the bitrate using the configuration parameter LambdaScaleTowardsNextQP. The sequences and target bitrates follow AHG17 (configurations of draft CfE).
In v2 the document added BDR comparisons with QP + 1 approach and fixed QP (QP closest to target bitrate) are included. The lambda scaling approach has more similar BDR compared to using fixed QP than to the QP + 1 rate matching approach.
JVET-AM0078: AHG17 Suggestion to enable DMVR encoder control for VTM anchor
This document is about the configuration parameter DMVREncMvSelect in VTM which will penalize the usage of DMVR when it is a risk that the tool produces visual artefacts. This typically occurs at lower bitrates when DMVR has more troubles to find the correct motion for all subblocks. In this contribution DMVREncMvSelect has been tested on the AHG17 test sequences and target bitrate’s and it is asserted that its enabling can improve visual quality in some cases. It is therefore suggested to enable DMVREncMvSelect for the VTM anchor in the context of subjective related tests, such as in the scope of AHG17. For rate matching the configuration parameter LambdaScaleTowardsNextQP has been used.
In v2 the document added BDR in comparisons with QP +1 rate matching and lambda scaling rate matching approach as in JVET-AM0045 are included. The BDR loss per class varies between -0.2% to 1.6% in comparison to the QP + 1 rate matching approach and between 0.7% to 1.4% in comparison to the lambda scaling rate matching approach.
JVET-AM0125: AHG17 Analysis on encoding time and mode cost test for draft CfE sequences
This document reports experimental results that show that runtime ratio of VTM high performance and ECM JVET-AL0245 are roughly 1.9x to 1.3x and 5.3x to 10.4x for R1 to R4, and encoding runtime and RDO count are correlated, and they vary greatly depending on sequences. For SDR_RA_HD, EncT per RDO count is 0.085, 0.080 msec/count in VTM 23.9 under default and high performance configurations respectively. On the other hand, it is 1.034 msec/count in ECM under JVET-AL0245 configuration, which is 13 times larger than that of VTM and it may reflect to 12x EncT in ECM.
JVET-AM0187: On Intra Period of Random Access Test Case for HD and below Content
This document reports that in today’s streaming applications, intra period is reportedly much larger than the 1 second period used in our CTC, especially when the adaptive intra picture inserting by rate control is excluded. It is proposed to use larger intra period of 2 seconds for content with resolution at or below HD, in random access test cases of CfP and future CTC to better align with practical applications while keeping a reasonable testing runtime. The benefit for UGC content in AHG17 is about 5% with 2s intra period.
JVET-AM0200: AHG17: Getting VTM to run five times faster
The document reports that the reduced runtime VTM encoder configuration #3 adopted at the 38th meeting doesn’t quite achieve the target of 0.2x encoder runtime under CfE test conditions. This contribution reports on software improvements that were recently made to VTM and proposes new encoder configurations to get closer to the target.
The document proposes to modify the encoder configuration variants for reduced runtime 2 (30%) and 3 (20%) present in VTM as follows:
- Reduced runtime configuration 2:
UseNonLinearAlfLuma : 0
SplitPredictAdaptMode : 2
MergeRdCandQuotaGpm : 5
ContentBasedFastQtbt : 1
AdaptBypassAffineMe : 1
MaxMTTHierarchyDepth : 2
MaxMTTHierarchyDepthISliceL : 2
MTS : 4
MTSIntraMaxCand : 3
MaxNumMergeCand : 5
AllowDisFracMMVD : 0
MaxMergeRdCandNumTotal : 5
- Reduced runtime configuration 3:
UseNonLinearAlfLuma : 0
SplitPredictAdaptMode : 2
MergeRdCandQuotaGpm : 4
MaxTTNonISlice : 32
MaxBTNonISlice : 64
ContentBasedFastQtbt : 1
AdaptBypassAffineMe : 1
CTUSize : 64
MaxMTTHierarchyDepth : 1
MaxMTTHierarchyDepthISliceL : 2
MaxMTTHierarchyDepthISliceC : 1
MTS : 4
MTSIntraMaxCand : 3
MaxNumMergeCand : 5
AllowDisFracMMVD : 0
AffineAmvr : 0
ISPFast : 1
FastMIP : 1
MaxMergeRdCandNumTotal : 4
AffineAmvrEncOpt : 0
JVET-AM0225: [AHG17] Encoder runtime for the constrained complexity configurations under the CfE draft test conditions
This document presents simulation results with the reliable encoder runtime for the constrained complexity encoder configurations of VTM 23.9 under the test conditions from the CfE draft. The simulation results show encoder runtime 65%, 53%, 25% with luma bd-rate 0.9%, 2.8%, 17.6% for the reduced runtime 1, 2 and 3 configurations correspondingly, relatively to the default RA configuration of VTM 23.9. HM-18.0 tested under the similar test conditions, shows 42% encoder runtime with 65.6% luma bd-rate relatively to the default configuration of VTM 23.9.
JVET-AM0237: On constrained encoding and decoding experiments
This document explores the trade-offs between coding efficiency and computational complexity for both encoder and decoder. During the previous meeting, multiple trade-off points were presented by changing ECM configurations. While certain configurations demonstrated encoder complexity levels approximately twice that of the VTM encoder, the corresponding decoder complexity failed to achieve a desired two times of VTM decoder. It was suggested that further modifications are necessary to effectively reduce decoder complexity alongside encoder complexity.
In the context of constrained encoding, specific configurations achieved encoder complexity levels approximately twice that of the VTM encoder while maintaining comparable BD-rate gains. However, the decoding time exhibited significant variations, ranging from x2, x4, to even x6 the decoding time of the VTM decoder.
This contribution emphasizes the importance of illustrating both encoding and decoding times to provide a comprehensive understanding of trade-offs within the constrained encoding category. It is suggested that decoding time information be discussed jointly with encoding time, and this is beneficial for CfE response and practical implementation of the next generation of video coding standard.
JVET-AM0238: Suggestion of GOP Size Setting of Random Access Configuration for Live-streaming Applications
This document states that latency is critical for live streaming applications, and a larger GOP size will lead to higher delay. Therefore, the GOP size in practical live streaming applications generally is much smaller than the value of 32 as specified in the current Common Test Conditions (CTC). To better reflect the use case, it is proposed to add one additional test configuration, setting the GOP size to be 4 in random access test cases of CfP and future CTC, ensuring better alignment with practical live streaming scenarios. In addition, the intra period is also adjusted to be 2 seconds instead of 1 second to better match typical live streaming settings in real applications.
Recommendations
The AHG recommended that the related input contributions be reviewed, and to further continue the study of encoding algorithm optimizations in JVET.